NVIDIA RTX 5090

Consumer GPU · Blackwell

Summary

The RTX 5090 is NVIDIA's current flagship gaming/retail GPU. With the addition of native FP4 support and 12GB more VRAM over the previous generation's RTX 4090, a single RTX 5090 is usually a great choice for inference of LLMs up to 30B parameters.

Launched
Q1 2025
VRAM
32 GB
Mem. bandwidth
1,792 GB/s
On-demand from
Compare prices from 5 providers →

Tech specs

VRAM 32 GB GDDR7
Memory bandwidth 1,792 GB/s
Interface PCIe
CUDA cores 21,760
Tensor cores 680 (Gen 5)
TDP 575 W
Supported data types
FP64FP32TF32FP16BF16FP8FP6FP4INT8

Cloud rental prices

5 providers · available in 6 countries · lowest price per GPU, per hour

Prices updated

On-demand
$0.48 – $0.99 /GPU/h
Reserved
$0.36 – $3.27 /GPU/h
Spot
$0.32 – $0.77 /GPU/h
Provider On-demand Reserved Spot
LeaderGPU $0.76
Nova $0.48 $0.36
Runpod $0.99
Sesterce $0.72
Vast.ai $0.67 $0.32
Compare all NVIDIA RTX 5090 cloud providers & configurations →

Models that fit in VRAM

Total VRAM From 16-bit inference 8-bit inference 4-bit inference
1× RTX 5090 32 GB $0.36/h Qwen3-4B-Instruct-2507Rio-3.0-Open-MiniNVIDIA-Nemotron-3-Nano-4B-BF16 gpt-oss-20bQwen3-4B-Instruct-2507Rio-3.0-Open-Mini GLM-4.7-FlashQwen3-Coder-30B-A3B-InstructNVIDIA-Nemotron-3-Nano-30B-A3B-BF16
2× RTX 5090 64 GB $0.72/h gpt-oss-20bQwen3-4B-Instruct-2507Rio-3.0-Open-Mini GLM-4.7-FlashQwen3-Coder-30B-A3B-InstructNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 GLM-4.5-AirQwen3-Coder-NextNVIDIA-Nemotron-3-Nano-30B-A3B-BF16
3× RTX 5090 96 GB $2.97/h GLM-4.7-FlashQwen3-Coder-30B-A3B-InstructNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 Qwen3-Coder-NextGLM-4.7-FlashNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 DeepSeek-V4-Flashgpt-oss-120bNVIDIA-Nemotron-3-Super-120B-A12B-BF16
4× RTX 5090 128 GB $3.31/h GLM-4.7-FlashQwen3-Coder-30B-A3B-InstructNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 GLM-4.5-AirQwen3-Coder-NextNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 DeepSeek-V4-Flashgpt-oss-120bNVIDIA-Nemotron-3-Super-120B-A12B-BF16
5× RTX 5090 160 GB $4.95/h GLM-4.7-FlashQwen3-Coder-30B-A3B-InstructNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 gpt-oss-120bNVIDIA-Nemotron-3-Super-120B-A12B-BF16GLM-4.5-Air MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b
6× RTX 5090 192 GB $5.94/h Qwen3-Coder-NextGLM-4.7-FlashNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 DeepSeek-V4-Flashgpt-oss-120bNVIDIA-Nemotron-3-Super-120B-A12B-BF16 MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b
7× RTX 5090 224 GB $6.93/h Qwen3-Coder-NextGLM-4.7-FlashNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 DeepSeek-V4-Flashgpt-oss-120bNVIDIA-Nemotron-3-Super-120B-A12B-BF16 MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b
8× RTX 5090 256 GB $2.88/h GLM-4.5-AirQwen3-Coder-NextNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b

Open-weights models that fit in VRAM. Estimated using 🤗 accelerate, plus approximation for up to 8K context.

Media

Peak theoretical performance

FP4 Tensor Core 1,676 TFLOPS
FP8 Tensor Core 838 TFLOPS
INT8 Tensor Core 838 TOPS
BF16 Tensor Core 419 TFLOPS
FP16 Tensor Core 419 TFLOPS
TF32 Tensor Core 104.8 TFLOPS
FP32 104.8 TFLOPS

Performance figures assume no sparsity; in cases where only sparse performance figures are published by the manufacturer, these are halved to give approximate dense performance.

Resources

Detailed documentation from the manufacturer.

Similar GPUs

Other accelerators you might compare.