NVIDIA H100 NVL

Data center GPU · Hopper

Summary

The highest-VRAM variant of NVIDIA's flagship Hopper-generation data center GPU, targeted at LLM inference.

For multi-GPU training with the H100, see the H100 SXM.

For an improved version of this GPU with more VRAM, see the H200 NVL.

Launched
Q1 2023
VRAM
94 GB
Mem. bandwidth
3,900 GB/s
On-demand from
Compare prices from 5 providers →

Tech specs

VRAM 94 GB HBM3
Memory bandwidth 3,900 GB/s
Interface PCIe
CUDA cores 14,592
Tensor cores 456 (Gen 4)
TDP 400 W
Supported data types
FP64FP32FP16BF16FP8INT8

Cloud rental prices

5 providers · available in 18 countries · lowest price per GPU, per hour

Prices updated

On-demand
$2.92 – $6.98 /GPU/h
Reserved
$3.58 – $5.24 /GPU/h
Spot
$0.81 – $1.40 /GPU/h
Provider On-demand Reserved Spot
Atlantic.Net $3.94 $3.58
Azure $6.98 $3.84 $1.40
Massed Compute $2.92
Runpod $3.19
Vast.ai $0.81
Compare all NVIDIA H100 NVL cloud providers & configurations →

Models that fit in VRAM

Total VRAM From 16-bit inference 8-bit inference 4-bit inference
1× H100 NVL 94 GB $3.11/h GLM-4.7-FlashQwen3-Coder-30B-A3B-InstructNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 Qwen3-Coder-NextGLM-4.7-FlashNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 DeepSeek-V4-Flashgpt-oss-120bNVIDIA-Nemotron-3-Super-120B-A12B-BF16
2× H100 NVL 188 GB $5.84/h Qwen3-Coder-NextGLM-4.7-FlashNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 DeepSeek-V4-Flashgpt-oss-120bNVIDIA-Nemotron-3-Super-120B-A12B-BF16 MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b
3× H100 NVL 282 GB $9.57/h gpt-oss-120bNVIDIA-Nemotron-3-Super-120B-A12B-BF16GLM-4.5-Air MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b
4× H100 NVL 376 GB $11.68/h DeepSeek-V4-Flashgpt-oss-120bNVIDIA-Nemotron-3-Super-120B-A12B-BF16 MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b
5× H100 NVL 470 GB $15.95/h DeepSeek-V4-Flashgpt-oss-120bNVIDIA-Nemotron-3-Super-120B-A12B-BF16 MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b DeepSeek-V3.2MiniMax-M2.7gpt-oss-120b
6× H100 NVL 564 GB $19.14/h MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b DeepSeek-V4-ProMiniMax-M2.7gpt-oss-120b
7× H100 NVL 658 GB $22.33/h MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b Kimi-K2-Instruct-0905DeepSeek-V4-ProMiniMax-M2.7
8× H100 NVL 752 GB $23.36/h MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b Kimi-K2-Instruct-0905DeepSeek-V4-ProMiniMax-M2.7

Open-weights models that fit in VRAM. Estimated using 🤗 accelerate, plus approximation for up to 8K context.

Media

Peak theoretical performance

FP8 Tensor Core 1,670 TFLOPS
INT8 Tensor Core 1,670 TOPS
BF16 Tensor Core 835 TFLOPS
FP16 Tensor Core 835 TFLOPS
TF32 Tensor Core 417 TFLOPS
FP32 60 TFLOPS
FP64 30 TFLOPS
FP64 Tensor Core 60 TFLOPS

Performance figures assume no sparsity; in cases where only sparse performance figures are published by the manufacturer, these are halved to give approximate dense performance.

Resources

Detailed documentation from the manufacturer.

Similar GPUs

Other accelerators you might compare.