$ whoami
Sanchit Gupta
AI Infrastructure Engineer @ Neuralix.ai
B.Tech CSE, IIIT Lucknow '25
$ cat what_i_do.txt
I keep large language models running on GPUs
that have no internet connection.
$ ./highlights.sh
β’ Core engineer on EKAM AI β an indigenous Defence
AI-as-a-Service platform under the MoD iDEX ADITI 2.0
initiative, launched at the Chanakya Defence Dialogue 2025
β’ Model serving on vLLM and NVIDIA Triton across GPU nodes
β continuous batching, KV-cache sizing, request scheduling
β’ Instrumented telemetry, profiled hot paths, tuned async
execution and caching β cut average API latency by 40%
β’ Delivery into secure air-gapped and VDI sites, where every
package and model weight has to be carried in offline$ cat architecture.txt
ββββββββββββββββββββββββββββββββββ
client ββββββββββββΆβ FastAPI gateway Β· React β
ββββββββββββββββββ¬ββββββββββββββββ
β
ββββββββββββββββββββββββββββΌβββββββββββββββββββββββββββ
βΌ βΌ βΌ
βββββββββββββββββββ βββββββββββββββββββ βββββββββββββββββββ
β vLLM β β NVIDIA Triton β β Qdrant β
β LLMs Β· SLMs β β OCR Β· VLM β β vector search β
β cont. batching β β reranking β β RAG retrieval β
ββββββββββ¬βββββββββ ββββββββββ¬βββββββββ βββββββββββββββββββ
ββββββββββββββ¬ββββββββββββββ
βΌ
βββββββββββββββββββ ββββββββββββββββββββββββββββ
β NVIDIA GPU βββββββββΆβ Prometheus Β· Grafana β
βββββββββββββββββββ ββββββββββββββββββββββββββββ
ββ no egress Β· no package mirror Β· no second chances ββ$ cat stack.txt
serving vLLM Β· NVIDIA Triton Β· Ollama Β· CUDA
models LLMs Β· SLMs Β· VLMs Β· OCR Β· embedding Β· reranking
retrieval Qdrant Β· RAG pipelines
backend Python Β· FastAPI Β· PostgreSQL Β· Redis Β· C++
platform Docker Β· Kubernetes Β· Linux Β· NGINX Β· Git
observe Prometheus Β· Grafana Β· Loki Β· Tempo Β· OpenTelemetry
frontend React Β· JavaScript
$ ls ~/projects
AI-inference/ self-hosted GPU platform on Kubernetes
url-shortener/ distributed, sub-100ms under load
AgroSmart/ soil and yield prediction, Django + ML
CampusConnect/ MERN admissions portal
$ echo $CONTACT
sanchitguptaghj@gmail.com