Skip to content

🛠️ Skills ​


AI Serving & GPU Infrastructure ​

Inference serversvLLM, NVIDIA Triton Inference Server, Ollama
OptimisationContinuous batching, KV-cache sizing, request scheduling, GPU memory tuning
ModelsLLMs, SLMs, VLMs, OCR, embedding and reranking models
LifecycleBenchmarking, validation, staging promotion, production rollout
RetrievalQdrant vector search, RAG pipelines, chunking and embedding

Backend ​

Python · FastAPI · REST APIs · async programming · microservices · C++ · C · SQL

Data: PostgreSQL · Redis · MongoDB

Platform & Deployment ​

Docker · Docker Compose · Kubernetes (MicroK8s) · Linux · NGINX · Git / GitHub

Specialist: air-gapped deployment · VDI environments · standalone application packaging · offline package mirroring

Observability ​

Prometheus · Grafana · Loki · Tempo · OpenTelemetry

Root-cause analysis across application, CUDA driver, network and hardware layers.

Frontend ​

React · JavaScript · HTML · CSS

Computer Science ​

Data structures and algorithms · system design · operating systems · computer networks · DBMS · concurrent programming · performance and memory optimisation


🏆 Problem solving

LeetCode Knight (1891) with 1000+ problems solved in C++ · Top 2000 in Google Code Jam 2022 · Top 2500 in Google Kick Start 2022 · Top 100 on GeeksforGeeks at IIIT Lucknow

<<>> with ♥️ by S@Nchit