NVIDIA L40
Data center GPU · Ada Lovelace
Summary
A capable Ada generation card that's a good fit for graphics/rendering workloads that make use of the RT cores.
For LLM inference, the NVIDIA L40s is usually a better fit, given its better FP8 performance. For multi-GPU training, look instead to the H100 or B200.
Launched
Q1 2023
VRAM
48 GB
Mem. bandwidth
864 GB/s
On-demand from
Tech specs
VRAM 48 GB GDDR6
Memory bandwidth 864 GB/s
Interface PCIe
CUDA cores 18,176
Tensor cores 568 (Gen 4)
TDP 300 W
Supported data types
FP64FP32TF32FP16BF16FP8INT8
Cloud rental prices
5 providers · available in 2 countries · lowest price per GPU, per hour
On-demand
$0.82 – $1.25 /GPU/h
Reserved
$0.70 /GPU/h
Spot
$0.78 – $0.80 /GPU/h
| Provider | On-demand | Reserved | Spot |
|---|---|---|---|
CoreWeave | $1.25 | — | $0.78 |
Hyperstack | $1.00 | $0.70 | $0.80 |
Massed Compute | $0.86 | — | — |
Runpod | $0.82 | — | — |
| | $0.97 | — | — |
Models that fit in VRAM
Open-weights models that fit in VRAM. Estimated using 🤗 accelerate, plus approximation for up to 8K context.
Media
Peak theoretical performance
FP8 Tensor Core 362 TFLOPS
INT8 Tensor Core 362 TOPS
INT4 Tensor Core 724 TOPS
BF16 Tensor Core 181 TFLOPS
FP16 Tensor Core 181 TFLOPS
TF32 Tensor Core 90.5 TFLOPS
FP32 90.5 TFLOPS
Performance figures assume no sparsity; in cases where only sparse performance figures are published by the manufacturer, these are halved to give approximate dense performance.
Resources
Detailed documentation from the manufacturer.



