NVIDIA RTX 4090

Consumer GPU · Ada Lovelace

Summary

The flagship gaming GPU of NVIDIA's Ada generation. Successor to the RTX 3090, and superseded by the RTX 5090.

Launched
Q4 2022
VRAM
24 GB
Mem. bandwidth
1,008 GB/s
On-demand from
Compare prices from 4 providers →

Tech specs

VRAM 24 GB GDDR6X
Memory bandwidth 1,008 GB/s
Interface PCIe
CUDA cores 16,384
Tensor cores 512 (Gen 4)
TDP 450 W
Supported data types
FP64FP32TF32FP16BF16FP8INT8

Cloud rental prices

4 providers · available in 2 countries · lowest price per GPU, per hour

Prices updated

On-demand
$0.66 – $2.05 /GPU/h
Reserved
$0.47 – $2.84 /GPU/h
Spot
$0.34 /GPU/h
Provider On-demand Reserved Spot
LeaderGPU $1.71 $0.47
Runpod $0.69
Sesterce $0.66
Vast.ai $0.34
Compare all NVIDIA RTX 4090 cloud providers & configurations →

Models that fit in VRAM

Total VRAM From 16-bit inference 8-bit inference 4-bit inference
1× RTX 4090 24 GB $0.47/h Qwen3-4B-Instruct-2507Rio-3.0-Open-MiniNVIDIA-Nemotron-3-Nano-4B-BF16 gpt-oss-20bQwen3-4B-Instruct-2507Rio-3.0-Open-Mini GLM-4.7-FlashQwen3-Coder-30B-A3B-InstructNVIDIA-Nemotron-3-Nano-30B-A3B-BF16
2× RTX 4090 48 GB $1.38/h gpt-oss-20bQwen3-4B-Instruct-2507Rio-3.0-Open-Mini GLM-4.7-FlashQwen3-Coder-30B-A3B-InstructNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 Qwen3-Coder-NextGLM-4.7-FlashNVIDIA-Nemotron-3-Nano-30B-A3B-BF16
3× RTX 4090 72 GB $2.07/h GLM-4.7-FlashQwen3-Coder-30B-A3B-InstructNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 GLM-4.7-FlashQwen3-Coder-30B-A3B-InstructNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 gpt-oss-120bNVIDIA-Nemotron-3-Super-120B-A12B-BF16GLM-4.5-Air
4× RTX 4090 96 GB $2.76/h GLM-4.7-FlashQwen3-Coder-30B-A3B-InstructNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 Qwen3-Coder-NextGLM-4.7-FlashNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 DeepSeek-V4-Flashgpt-oss-120bNVIDIA-Nemotron-3-Super-120B-A12B-BF16
5× RTX 4090 120 GB $3.45/h GLM-4.7-FlashQwen3-Coder-30B-A3B-InstructNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 GLM-4.5-AirQwen3-Coder-NextNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 DeepSeek-V4-Flashgpt-oss-120bNVIDIA-Nemotron-3-Super-120B-A12B-BF16
6× RTX 4090 144 GB $4.14/h GLM-4.7-FlashQwen3-Coder-30B-A3B-InstructNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 gpt-oss-120bNVIDIA-Nemotron-3-Super-120B-A12B-BF16GLM-4.5-Air MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b
7× RTX 4090 168 GB $4.83/h GLM-4.7-FlashQwen3-Coder-30B-A3B-InstructNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 gpt-oss-120bNVIDIA-Nemotron-3-Super-120B-A12B-BF16GLM-4.5-Air MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b
8× RTX 4090 192 GB $5.52/h Qwen3-Coder-NextGLM-4.7-FlashNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 DeepSeek-V4-Flashgpt-oss-120bNVIDIA-Nemotron-3-Super-120B-A12B-BF16 MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b

Open-weights models that fit in VRAM. Estimated using 🤗 accelerate, plus approximation for up to 8K context.

Peak theoretical performance

FP8 Tensor Core 660.6 TFLOPS
INT8 Tensor Core 660.6 TOPS
BF16 Tensor Core 330.3 TFLOPS
FP16 Tensor Core 330.3 TFLOPS
TF32 Tensor Core 82.6 TFLOPS
FP32 82.6 TFLOPS

Performance figures assume no sparsity; in cases where only sparse performance figures are published by the manufacturer, these are halved to give approximate dense performance.

Resources

Detailed documentation from the manufacturer.

Similar GPUs

Other accelerators you might compare.