Inference Engineer · Sarvam

A GPU whisperer. Making models fast, cheap, and reliable.

I work across the inference stack, from CUDA kernels to Kubernetes autoscaling. This is my field notebook: a day-by-day journey across the stack, every post built, measured, and verified by measurement.

Your whole life, in tokens

drag the sliders

The journey

the journey

Latest posts

built by hand