Runtime
Going deeperMaking one model fast on one GPU-backed instance.
Sr. Cloud Infrastructure Engineer at Simplismart
I build the infrastructure AI inference runs on.
GPUs, inference runtimes, Kubernetes, and the clusters and clouds underneath them. I like understanding what happens below the abstractions.
About
I'm an infrastructure engineer working my way deeper into AI systems.
At Simplismart I work on the platform that serves models in production. That means Kubernetes clusters full of NVIDIA GPUs, spread across AWS, Azure, GCP and on-prem hardware, plus the tooling to deploy, scale and watch them.
Outside work I keep going further down the stack, into inference runtimes, CUDA and how a GPU actually spends its clock cycles. I learn by building small experiments, breaking things, and writing up what I find.
I got here the long way round. I started out building Flutter apps and ML models at hackathons, then spent two years at Baton Systems, where Java integration work slowly turned into Kubernetes work. Since October 2025 it has been infrastructure full time.
How I think
Doing inference well means getting three things right. My day job lives mostly in the second and third. The first is where I'm going deeper.
Making one model fast on one GPU-backed instance.
Scaling that across GPUs, nodes, clusters, regions and clouds, while keeping latency, availability and utilization where they need to be.
Giving engineers the right abstraction to deploy and run models, without taking away control.
Languages · Python, Go, C/C++, Java, Bash
Experience
Simplismart · Bengaluru · Apr 2026 – now
Simplismart runs a platform for fast, production AI inference. I work on the infrastructure underneath it.
Simplismart · Oct 2025 – Mar 2026
Baton Systems · Chennai
Ernst & Young (EY) · Kochi
Christ College of Engineering · Irinjalakuda, Kerala · CGPA 8.63
Case study · Simplismart
Separate Kubernetes clusters are easy to run until you need them to behave like one platform. I implemented federation with Karmada, so workloads across independent clusters are managed from a single place.
GPU capacity sits in different clouds and regions. Separate clusters also keep failures contained and tenants apart.
You declare a workload once. Propagation policies decide which member clusters run it, and override policies handle the per-cluster differences.
One more control plane to operate, and debugging that now crosses cluster boundaries.
Notebook
What I'm studying right now. Findings go out on X first, and longer write-ups land on Substack.
How the scheduler mixes prefill and decode, how PagedAttention hands out KV-cache blocks, and when speculative decoding with a draft model actually pays off.
SMs, warps, Tensor Cores, HBM and L2 on the H100. Next up is small CUDA experiments that measure occupancy, L2 hit rate and the point where a kernel turns memory-bound.
Routing on queue depth and KV-cache locality instead of round robin, and falling back to a managed provider when a pool runs hot.
The pod lifecycle, scheduling, device plugins, admission webhooks, kube-proxy and DNS. The parts you usually meet for the first time during an incident.
A clock is a metronome. Each tick captures state, and the transistors do the actual work between ticks.
GitHub
Before infrastructure I built apps and ML models, mostly at hackathons. A few from GitHub that still hold up.
A bank simulator that updates cash and non-cash accounts from SWIFT messages. Spring Boot microservices, tuned for concurrency with ExecutorService, handling around 1,000 TPS.
An AI travel guide. An on-device TFLite model recognises monuments and narrates their history, with voice-to-voice translation for travellers. Won Hack@Arch 2022.
Field-level identification of pests and plant diseases, with a TFLite image classifier running inside a Flutter app.
Wolfram's Rule 30 cellular automaton rendered with Ebiten, to see how a fully deterministic rule produces output that looks random.
Small things built out of curiosity: Bloom filters, why i++ breaks under concurrency, Java annotations and functional interfaces.
Sends WhatsApp messages to a list of saved or unsaved contacts from a spreadsheet, driven by Selenium.
Outside work
I like being around India's open-source and cloud-native communities, from IndiaFOSS to LibreMinds. And I write in public, mostly on X.
Community work on Cloud Native Summit Kerala 2026 and the vLLM Inference Meetup.
A talk on eBPF.
A talk on browser automation with Selenium.
One international and four national, including Hac'KP 2021 by Kerala Police (dark-web monitoring), Hack@Arch 2022 (AI travel guide), RIBC 2022 (decentralised cab booking) and EY WhyHack 2022 (pothole detection).
Contact