NVIDIA L4

Data center GPU · Ada Lovelace

Summary

The NVIDIA's L4 is a low-power, efficient GPU which works well for video transcoding workloads. It is also capable of inference on very small LLMs, but the larger and more powerful NVIDIA L40s is generally a better fit for inference workloads.

Launched
Q2 2023
VRAM
24 GB
Mem. bandwidth
300 GB/s
On-demand from
Compare prices from 8 providers →

Tech specs

VRAM 24 GB GDDR6
Memory bandwidth 300 GB/s
Interface PCIe
CUDA cores 7,424
Tensor cores 240 (Gen 4)
TDP 72 W
Supported data types
FP64FP32TF32FP16BF16FP8INT8

Cloud rental prices

8 providers · available in 18 countries · lowest price per GPU, per hour

Prices updated

On-demand
$0.44 – $1.67 /GPU/h
Reserved
$0.37 – $1.09 /GPU/h
Spot
$0.29 /GPU/h
Provider On-demand Reserved Spot
AceCloud ? $0.60
AWS ? $0.80 $0.37
Jarvislabs $0.44 $0.40 $0.29
OVHcloud ? $0.87
Runpod $0.49
Scaleway ? $0.91
Seeweb ? $0.44 $0.37
Sesterce $1.05
Compare all NVIDIA L4 cloud providers & configurations →

Models that fit in VRAM

Total VRAM From 16-bit inference 8-bit inference 4-bit inference
1× L4 24 GB $0.37/h Qwen3-4B-Instruct-2507Rio-3.0-Open-MiniNVIDIA-Nemotron-3-Nano-4B-BF16 gpt-oss-20bQwen3-4B-Instruct-2507Rio-3.0-Open-Mini GLM-4.7-FlashQwen3-Coder-30B-A3B-InstructNVIDIA-Nemotron-3-Nano-30B-A3B-BF16
2× L4 48 GB $0.76/h gpt-oss-20bQwen3-4B-Instruct-2507Rio-3.0-Open-Mini GLM-4.7-FlashQwen3-Coder-30B-A3B-InstructNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 Qwen3-Coder-NextGLM-4.7-FlashNVIDIA-Nemotron-3-Nano-30B-A3B-BF16
3× L4 72 GB $1.47/h GLM-4.7-FlashQwen3-Coder-30B-A3B-InstructNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 GLM-4.7-FlashQwen3-Coder-30B-A3B-InstructNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 gpt-oss-120bNVIDIA-Nemotron-3-Super-120B-A12B-BF16GLM-4.5-Air
4× L4 96 GB $1.50/h GLM-4.7-FlashQwen3-Coder-30B-A3B-InstructNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 Qwen3-Coder-NextGLM-4.7-FlashNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 DeepSeek-V4-Flashgpt-oss-120bNVIDIA-Nemotron-3-Super-120B-A12B-BF16
5× L4 120 GB $2.45/h GLM-4.7-FlashQwen3-Coder-30B-A3B-InstructNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 GLM-4.5-AirQwen3-Coder-NextNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 DeepSeek-V4-Flashgpt-oss-120bNVIDIA-Nemotron-3-Super-120B-A12B-BF16
6× L4 144 GB $2.94/h GLM-4.7-FlashQwen3-Coder-30B-A3B-InstructNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 gpt-oss-120bNVIDIA-Nemotron-3-Super-120B-A12B-BF16GLM-4.5-Air MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b
7× L4 168 GB $3.43/h GLM-4.7-FlashQwen3-Coder-30B-A3B-InstructNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 gpt-oss-120bNVIDIA-Nemotron-3-Super-120B-A12B-BF16GLM-4.5-Air MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b
8× L4 192 GB $3.00/h Qwen3-Coder-NextGLM-4.7-FlashNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 DeepSeek-V4-Flashgpt-oss-120bNVIDIA-Nemotron-3-Super-120B-A12B-BF16 MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b
9× L4 216 GB $4.41/h Qwen3-Coder-NextGLM-4.7-FlashNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 DeepSeek-V4-Flashgpt-oss-120bNVIDIA-Nemotron-3-Super-120B-A12B-BF16 MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b

Open-weights models that fit in VRAM. Estimated using 🤗 accelerate, plus approximation for up to 8K context.

Media

Peak theoretical performance

FP8 Tensor Core 242 TFLOPS
INT8 Tensor Core 242 TOPS
BF16 Tensor Core 121 TFLOPS
FP16 Tensor Core 121 TFLOPS
TF32 Tensor Core 60 TFLOPS
FP32 30.3 TFLOPS

Performance figures assume no sparsity; in cases where only sparse performance figures are published by the manufacturer, these are halved to give approximate dense performance.

Resources

Detailed documentation from the manufacturer.