Sr. Cloud Infrastructure Engineer at Simplismart

AdityaKrishnan

I build the infrastructure AI inference runs on.

GPUs, inference runtimes, Kubernetes, and the clusters and clouds underneath them. I like understanding what happens below the abstractions.

Bengaluru, IN · Kubernetes · vLLM · CUDA · GPU systems

POST/v1/chat/completions 200 OK
Prompt
What does the KV cache do?
Tokens
It keeps the keys and values of tokens already processed, so each new token only computes its own.
  1. Cloud AWS · Azure · GCP · on-premGPUs wherever there is capacity
  2. Fleet Karmadaone control plane, many clusters
  3. Cluster Kubernetes · KEDA · KAI-Schedulerautoscale on queue depth
  4. Node NVIDIA device plugin · MIGone H100, several workloads
  5. Runtime vLLM · SGLang · TensorRT-LLMcontinuous batching · PagedAttention
  6. Kernel CUDAwarps of 32 threads
  7. Silicon H100 · B200 · L40SH100 SXM · 132 SMs · 80 GB HBM3
Fig. 1What sits under a single inference request, cloud to silicon. These are the layers I work in.

About

Now

I'm an infrastructure engineer working my way deeper into AI systems.

At Simplismart I work on the platform that serves models in production. That means Kubernetes clusters full of NVIDIA GPUs, spread across AWS, Azure, GCP and on-prem hardware, plus the tooling to deploy, scale and watch them.

Outside work I keep going further down the stack, into inference runtimes, CUDA and how a GPU actually spends its clock cycles. I learn by building small experiments, breaking things, and writing up what I find.

I got here the long way round. I started out building Flutter apps and ML models at hackathons, then spent two years at Baton Systems, where Java integration work slowly turned into Kubernetes work. Since October 2025 it has been infrastructure full time.

Portrait of Aditya Krishnan
Based in
Bengaluru, India
Working on
Inference infrastructure
Experience
3+ years in production
Hardware
NVIDIA H100, B200, L40S
Clouds
AWS, Azure, GCP, on-prem
Studied
B.Tech, Computer Science

How I think

Inference takes three layers

Doing inference well means getting three things right. My day job lives mostly in the second and third. The first is where I'm going deeper.

Runtime

Going deeper

Making one model fast on one GPU-backed instance.

  • vLLM
  • SGLang
  • TensorRT-LLM
  • NVIDIA Dynamo
  • continuous batching
  • PagedAttention
  • KV cache
  • speculative decoding
  • FP8 / INT4
  • CUDA

Infrastructure

Day to day

Scaling that across GPUs, nodes, clusters, regions and clouds, while keeping latency, availability and utilization where they need to be.

  • Kubernetes
  • kubeadm
  • Karmada
  • KEDA
  • KAI-Scheduler
  • MIG
  • NVIDIA GDS
  • WekaFS
  • WireGuard
  • multi-tenancy
  • AWS
  • Azure
  • GCP
  • Terraform

Tooling

Day to day

Giving engineers the right abstraction to deploy and run models, without taking away control.

  • Helm
  • ArgoCD
  • Prometheus
  • Grafana Cloud
  • Loki
  • Mimir
  • NVIDIA DCGM
  • ClickHouse
  • Kafka
  • PagerDuty
  • incident response
  • RCA

Languages · Python, Go, C/C++, Java, Bash

Experience

Work

lower GPU pod cold-start latency, from warm-pool and scale-up work across three clouds
90%
workloads sharing a single H100, with MIG and KAI-Scheduler
4+
less metrics cardinality in a zero-downtime move to Grafana Cloud
up to80%
from Cloud Infrastructure Engineer to Senior, Oct 2025 to Apr 2026
6 mo
Oct 2025 – nowRunning

Sr. Cloud Infrastructure Engineer

Simplismart · Bengaluru · Apr 2026 – now

Simplismart runs a platform for fast, production AI inference. I work on the infrastructure underneath it.

  • Built a self-hosted Kubernetes platform end to end with kubeadm for a multi-tenant GPU inference fleet.
  • Rolled out KAI-Scheduler with MIG-based fractional GPU sharing, priority-based node affinity and NVIDIA GDS, so 4+ workloads can share a single H100.
  • Led a zero-downtime migration from self-hosted Loki and Mimir to Grafana Cloud, cutting metrics cardinality by up to 80% and scoping alerts per org for customer-facing SLAs.
  • Implemented multi-cluster federation with Karmada, bringing independent clusters under one control plane.
  • Drove warm-pool stabilization and faster scale-up across AWS, Azure and GCP, cutting GPU pod cold-start latency by 90% for high-concurrency inference.

Cloud Infrastructure Engineer

Simplismart · Oct 2025 – Mar 2026

  • Hardened production clusters with RBAC, WAF, ZeroSSL/TLS automation and multi-cloud connectivity over WireGuard VPC peering, closing security gaps ahead of customer onboarding.
  • Ended recurring storage outages by fixing WekaFS and Loki/Mimir PVC issues, and enabled RWX volumes on all three clouds for shared model storage.
  • Cut GPU idle-compute cost with KEDA-driven autoscaling and Cluster Autoscaler node groups.
  • Kubernetes
  • kubeadm
  • KAI-Scheduler
  • MIG
  • Karmada
  • KEDA
  • ArgoCD
  • Helm
  • Terraform
  • Grafana Cloud
  • WekaFS
Aug 2023 – Sep 2025Completed

Software Engineer II, Infrastructure

Baton Systems · Chennai

  • Set up HPA, node affinity and pod tolerations in Kubernetes to absorb traffic spikes while keeping node utilization efficient.
  • Built a Kubernetes-native, low-code integration tool on Camel Karavan and AtlasMap, so non-engineers could build ETL workflows.
  • Built an SFTP health check wired into PagerDuty that cut manual intervention by 30%.
  • Wrote an ingestion automation tool that took environment setup from 5 days to 5 hours.
  • Earned 3 Spot Awards (Q1 2024, Q3 2024, Q1 2025) for automation and infrastructure work.
  • Java
  • Spring Boot
  • Kubernetes
  • Helm
  • Apache Camel
  • MySQL
  • PagerDuty
Oct 2022 – Feb 2023Completed

Software Engineer Intern

Ernst & Young (EY) · Kochi

  • Built a fleet optimization model in Python using simulated annealing and meta-heuristics, to cut logistics costs across changing pick-up and drop points.
  • Python
  • Docker
  • optimization
2019 – 2023Completed

B.Tech, Computer Science & Engineering

Christ College of Engineering · Irinjalakuda, Kerala · CGPA 8.63

Case study · Simplismart

One control plane for many clusters

Separate Kubernetes clusters are easy to run until you need them to behave like one platform. I implemented federation with Karmada, so workloads across independent clusters are managed from a single place.

Why more than one cluster

GPU capacity sits in different clouds and regions. Separate clusters also keep failures contained and tenants apart.

What Karmada adds

You declare a workload once. Propagation policies decide which member clusters run it, and override policies handle the per-cluster differences.

What it costs

One more control plane to operate, and debugging that now crosses cluster boundaries.

Notebook

On the bench

What I'm studying right now. Findings go out on X first, and longer write-ups land on Substack.

runtimevLLM internals

How the scheduler mixes prefill and decode, how PagedAttention hands out KV-cache blocks, and when speculative decoding with a draft model actually pays off.

siliconGPUs below the framework

SMs, warps, Tensor Cores, HBM and L2 on the H100. Next up is small CUDA experiments that measure occupancy, L2 hit rate and the point where a kernel turns memory-bound.

clusterRouting inference traffic

Routing on queue depth and KV-cache locality instead of round robin, and falling back to a managed provider when a pool runs hot.

nodeKubernetes internals

The pod lifecycle, scheduling, device plugins, admission webhooks, kube-proxy and DNS. The parts you usually meet for the first time during an incident.

A clock is a metronome. Each tick captures state, and the transistors do the actual work between ticks.
From my notes on GPU clocks

GitHub

Things I've built

Before infrastructure I built apps and ML models, mostly at hackathons. A few from GitHub that still hold up.

SwiftPay

★ 2

A bank simulator that updates cash and non-cash accounts from SWIFT messages. Spring Boot microservices, tuned for concurrency with ExecutorService, handling around 1,000 TPS.

  • Java
  • Spring Boot
  • Docker
  • AWS

TraWell

★ 25

An AI travel guide. An on-device TFLite model recognises monuments and narrates their history, with voice-to-voice translation for travellers. Won Hack@Arch 2022.

  • Flutter
  • TFLite
  • Spring Boot
Site ↗

AgroMed

★ 11

Field-level identification of pests and plant diseases, with a TFLite image classifier running inside a Flutter app.

  • Flutter
  • TFLite
  • Firebase

Wolfram's Rule 30 cellular automaton rendered with Ebiten, to see how a fully deterministic rule produces output that looks random.

  • Go
  • Ebiten

Small things built out of curiosity: Bloom filters, why i++ breaks under concurrency, Java annotations and functional interfaces.

  • Java
  • concurrency

Sends WhatsApp messages to a list of saved or unsaved contacts from a spreadsheet, driven by Selenium.

  • Python
  • Selenium
All repositories on GitHub

Outside work

Community & writing

I like being around India's open-source and cloud-native communities, from IndiaFOSS to LibreMinds. And I write in public, mostly on X.

Talks & community

  • 2026

    LibreMinds

    Community work on Cloud Native Summit Kerala 2026 and the vLLM Inference Meetup.

  • 2025

    Speaker, DevDay Chennai

    A talk on eBPF.

  • 2025

    Speaker, Kochi FOSS

    A talk on browser automation with Selenium.

  • College

    Five hackathon wins

    One international and four national, including Hac'KP 2021 by Kerala Police (dark-web monitoring), Hack@Arch 2022 (AI travel guide), RIBC 2022 (decentralised cab booking) and EY WhyHack 2022 (pothole detection).

  • 2022

    Open source

    Patches to EvaMaria and Userge, two Python Telegram bots. My EvaMaria fork has been forked more than 120 times.

Contact

Working on inference, GPUs or Kubernetes? Let's talk.