NVIDIA L40S

Data center GPU · Ada Lovelace

Summary

NVIDIA's L40s FP8 performance makes it a great inference workhorse. For graphics/rendering workloads, consider the similar (but usually cheaper) NVIDIA L40. For multi-GPU training, look instead to the H100 or B200.

Launched
Q3 2023
VRAM
48 GB
Mem. bandwidth
864 GB/s
On-demand from
Compare prices from 17 providers →

Tech specs

VRAM 48 GB GDDR6
Memory bandwidth 864 GB/s
Interface PCIe
CUDA cores 18,176
Tensor cores 568 (Gen 4)
TDP 350 W
Supported data types
FP64FP32TF32FP16BF16FP8INT8

Cloud rental prices

17 providers · available in 14 countries · lowest price per GPU, per hour

Prices updated

On-demand
$0.88 – $3.77 /GPU/h
Reserved
$0.78 – $2.37 /GPU/h
Spot
$0.48 – $0.99 /GPU/h
Compare all NVIDIA L40S cloud providers & configurations →

Models that fit in VRAM

Total VRAM From 16-bit inference 8-bit inference 4-bit inference
1× L40S 48 GB $0.80/h gpt-oss-20bQwen3-4B-Instruct-2507Rio-3.0-Open-Mini GLM-4.7-FlashQwen3-Coder-30B-A3B-InstructNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 Qwen3-Coder-NextGLM-4.7-FlashNVIDIA-Nemotron-3-Nano-30B-A3B-BF16
2× L40S 96 GB $1.64/h GLM-4.7-FlashQwen3-Coder-30B-A3B-InstructNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 Qwen3-Coder-NextGLM-4.7-FlashNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 DeepSeek-V4-Flashgpt-oss-120bNVIDIA-Nemotron-3-Super-120B-A12B-BF16
3× L40S 144 GB $2.97/h GLM-4.7-FlashQwen3-Coder-30B-A3B-InstructNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 gpt-oss-120bNVIDIA-Nemotron-3-Super-120B-A12B-BF16GLM-4.5-Air MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b
4× L40S 192 GB $3.28/h Qwen3-Coder-NextGLM-4.7-FlashNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 DeepSeek-V4-Flashgpt-oss-120bNVIDIA-Nemotron-3-Super-120B-A12B-BF16 MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b
5× L40S 240 GB $4.95/h GLM-4.5-AirQwen3-Coder-NextNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 DeepSeek-V4-Flashgpt-oss-120bNVIDIA-Nemotron-3-Super-120B-A12B-BF16 MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b
6× L40S 288 GB $5.94/h gpt-oss-120bNVIDIA-Nemotron-3-Super-120B-A12B-BF16GLM-4.5-Air MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b
7× L40S 336 GB $6.93/h gpt-oss-120bNVIDIA-Nemotron-3-Super-120B-A12B-BF16GLM-4.5-Air MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b
8× L40S 384 GB $6.22/h DeepSeek-V4-Flashgpt-oss-120bNVIDIA-Nemotron-3-Super-120B-A12B-BF16 MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b
10× L40S 480 GB $15.00/h DeepSeek-V4-Flashgpt-oss-120bNVIDIA-Nemotron-3-Super-120B-A12B-BF16 MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b DeepSeek-V4-ProMiniMax-M2.7gpt-oss-120b

Open-weights models that fit in VRAM. Estimated using 🤗 accelerate, plus approximation for up to 8K context.

Media

Peak theoretical performance

FP8 Tensor Core 733 TFLOPS
INT8 Tensor Core 733 TOPS
INT4 Tensor Core 733 TOPS
BF16 Tensor Core 362 TFLOPS
FP16 Tensor Core 362 TFLOPS
TF32 Tensor Core 183 TFLOPS
FP32 91.6 TFLOPS

Performance figures assume no sparsity; in cases where only sparse performance figures are published by the manufacturer, these are halved to give approximate dense performance.

Resources

Detailed documentation from the manufacturer.