L40S is the Ada Lovelace card that bridges the gap between the consumer RTX line and the datacenter Hopper line. It's the "AI but not H100" option, and it's surprisingly capable.
What changed from L4
- More memory. 48GB GDDR6, up from L4's 24GB. And 864GB/s bandwidth, nearly 3x L4's 300GB/s.
- Full Ada die. 142 SMs, up from L4's 58. This is the full AD102 die, not a cut-down part.
- FP8 tensor cores. 733 TFLOPS FP8, the headline number. Ada's 4th-gen tensor cores with FP8 support.
- No NVLink. L40S is PCIe Gen4 x16 only, like L4 and T4. Multi-GPU parallelism goes over PCIe (64 GB/s), making tensor parallelism impractical beyond 2 GPUs.
- Higher power. 350W TDP, up from L4's 72W. It's a real datacenter card now.
The numbers
- FP8: 733 TFLOPS dense.
- FP16: 366 TFLOPS dense.
- Memory: 48GB GDDR6, 864GB/s.
- Power: 350W TDP.
Why it matters for inference
L40S is the "good enough for a 13B-30B model" card. 48GB fits a 70B model in INT4 (GPTQ/AWQ) on a single GPU, no tensor parallelism needed. Its FP8 support is real but not as complete as Hopper's Transformer Engine, so the FP8 win is smaller.
The interesting niche: L40S is often the cheapest way to get 48GB of GPU memory, making it a favorite for fine-tuning and for serving mid-size models where you need more memory than an L4 but don't want to pay H100 prices. Deploy it as independent replicas, not parallel-GPU training.
Sources
- NVIDIA L40S product page: official specs.
- NVIDIA Ada Lovelace architecture: the family L40S belongs to.