Inference Engineer · Sarvam
A GPU whisperer. Making models fast, cheap, and reliable.
I work across the inference stack, from CUDA kernels to Kubernetes autoscaling. This is my field notebook: a day-by-day journey across the stack, every post built, measured, and verified by measurement.