Manual · GPU

L4: Ada's answer to T4

24GB, AV1 encode, RTX IO, and the modern edge card. L4 is what happens when Ada Lovelace meets the datacenter.

L4 is T4's successor, built on the Ada Lovelace architecture. It's the modern answer to the question T4 answered: how do you make a small, efficient GPU that's still useful for AI?

What changed from T4

The numbers

Why it matters for inference

L4 is the modern edge card. Its 24GB fits a 7B model in FP16 (14GB) or a 13B model in INT8 (13GB), with room for KV cache. The AV1 encoders make it the default for video AI, and the low power means it slots into existing servers without power upgrades.

For inference, L4 is the "good enough, cheap" option: not a B200, but for a 7B model serving high volume, it's often the most cost-effective per token.

Sources

Back to the manual