🛠️ Skills
AI Serving & GPU Infrastructure
| Inference servers | vLLM, NVIDIA Triton Inference Server, Ollama |
| Optimisation | Continuous batching, KV-cache sizing, request scheduling, GPU memory tuning |
| Models | LLMs, SLMs, VLMs, OCR, embedding and reranking models |
| Lifecycle | Benchmarking, validation, staging promotion, production rollout |
| Retrieval | Qdrant vector search, RAG pipelines, chunking and embedding |
Backend
Python · FastAPI · REST APIs · async programming · microservices · C++ · C · SQL
Data: PostgreSQL · Redis · MongoDB
Platform & Deployment
Docker · Docker Compose · Kubernetes (MicroK8s) · Linux · NGINX · Git / GitHub
Specialist: air-gapped deployment · VDI environments · standalone application packaging · offline package mirroring
Observability
Prometheus · Grafana · Loki · Tempo · OpenTelemetry
Root-cause analysis across application, CUDA driver, network and hardware layers.
Frontend
React · JavaScript · HTML · CSS
Computer Science
Data structures and algorithms · system design · operating systems · computer networks · DBMS · concurrent programming · performance and memory optimisation
🏆 Problem solving
LeetCode Knight (1891) with 1000+ problems solved in C++ · Top 2000 in Google Code Jam 2022 · Top 2500 in Google Kick Start 2022 · Top 100 on GeeksforGeeks at IIIT Lucknow
