Skip to content

🛠️ Skills


AI Serving & GPU Infrastructure

Inference serversvLLM, NVIDIA Triton Inference Server, Ollama
OptimisationContinuous batching, KV-cache sizing, request scheduling, GPU memory tuning
ModelsLLMs, SLMs, VLMs, OCR, embedding and reranking models
LifecycleBenchmarking, validation, staging promotion, production rollout
RetrievalQdrant vector search, RAG pipelines, chunking and embedding

Backend

Python · FastAPI · REST APIs · async programming · microservices · C++ · C · SQL

Data: PostgreSQL · Redis · MongoDB

Platform & Deployment

Docker · Docker Compose · Kubernetes (MicroK8s) · Linux · NGINX · Git / GitHub

Specialist: air-gapped deployment · VDI environments · standalone application packaging · offline package mirroring

Observability

Prometheus · Grafana · Loki · Tempo · OpenTelemetry

Root-cause analysis across application, CUDA driver, network and hardware layers.

Frontend

React · JavaScript · HTML · CSS

Computer Science

Data structures and algorithms · system design · operating systems · computer networks · DBMS · concurrent programming · performance and memory optimisation


🏆 Problem solving

LeetCode Knight (1891) with 1000+ problems solved in C++ · Top 2000 in Google Code Jam 2022 · Top 2500 in Google Kick Start 2022 · Top 100 on GeeksforGeeks at IIIT Lucknow

<<>> with ♥️ by S@Nchit