NVIDIA H100

Data center GPU · Hopper · SXM, PCIe, NVL variants

Summary

NVIDIA's flagship Hopper-generation data center GPU. The H100's three variants share the same core Hopper architecture but differ in performance, form factor, interconnect, memory, and power.

The H100 SXM is the highest-performance variant, and the most widely available. It supports fast NVLink connectivity between GPUs, making it the default for multi-GPU training. The H100 PCIe is a lower-power card in a PCIe form factor. It is well suited to inference and single- or few-GPU training jobs. The H100 NVL has more memory (94 GB), alongside higher memory bandwidth, and is targeted at LLM inference.

Launched
Q4 2022
VRAM
80 - 94 GB
Mem. bandwidth
2,000 - 3,900 GB/s
On-demand from
Compare prices from 28 providers →

Tech specs

Spec H100 SXMH100 PCIeH100 NVL
VRAM 80 GB HBM380 GB HBM2e94 GB HBM3
Memory bandwidth 3,350 GB/s2,000 GB/s3,900 GB/s
Interface SXM PCIe PCIe
CUDA cores 16,896 14,592 14,592
Tensor cores 528 (Gen 4)456 (Gen 4)456 (Gen 4)
TDP 700 W350 W400 W
Supported data types
FP64FP32FP16BF16FP8INT8

Cloud rental prices

28 providers · available in 31 countries · lowest price per GPU, per hour

Prices updated

On-demand
$2.50 – $3.39 /GPU/h
Reserved
$1.18 – $3.88 /GPU/h
Compare all H100 PCIe cloud providers & configurations →
On-demand
$2.92 – $6.98 /GPU/h
Reserved
$3.58 – $5.24 /GPU/h
Spot
$0.81 – $1.40 /GPU/h
Provider On-demand Reserved Spot
Atlantic.Net $3.94 $3.58
Azure $6.98 $3.84 $1.40
Massed Compute $2.92
Runpod $3.19
Vast.ai $0.81
Compare all H100 NVL cloud providers & configurations →

Models that fit in VRAM

Total VRAM From 16-bit inference 8-bit inference 4-bit inference
1× H100 80 GB $1.18/h GLM-4.7-FlashQwen3-Coder-30B-A3B-InstructNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 GLM-4.7-FlashQwen3-Coder-30B-A3B-InstructNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 gpt-oss-120bNVIDIA-Nemotron-3-Super-120B-A12B-BF16GLM-4.5-Air
2× H100 160 GB $3.50/h GLM-4.7-FlashQwen3-Coder-30B-A3B-InstructNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 gpt-oss-120bNVIDIA-Nemotron-3-Super-120B-A12B-BF16GLM-4.5-Air MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b
3× H100 240 GB $8.67/h GLM-4.5-AirQwen3-Coder-NextNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 DeepSeek-V4-Flashgpt-oss-120bNVIDIA-Nemotron-3-Super-120B-A12B-BF16 MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b
4× H100 320 GB $7.00/h gpt-oss-120bNVIDIA-Nemotron-3-Super-120B-A12B-BF16GLM-4.5-Air MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b
5× H100 400 GB $14.45/h DeepSeek-V4-Flashgpt-oss-120bNVIDIA-Nemotron-3-Super-120B-A12B-BF16 MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b DeepSeek-V3.2MiniMax-M2.7gpt-oss-120b
6× H100 480 GB $17.34/h DeepSeek-V4-Flashgpt-oss-120bNVIDIA-Nemotron-3-Super-120B-A12B-BF16 MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b DeepSeek-V4-ProMiniMax-M2.7gpt-oss-120b
7× H100 560 GB $15.50/h MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b DeepSeek-V4-ProMiniMax-M2.7gpt-oss-120b
8× H100 640 GB $14.00/h MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b Kimi-K2-Instruct-0905DeepSeek-V4-ProMiniMax-M2.7

Open-weights models that fit in VRAM. Estimated using 🤗 accelerate, plus approximation for up to 8K context.

Media

Peak theoretical performance

Spec H100 SXMH100 PCIeH100 NVL
FP8 Tensor Core 1,979 TFLOPS1,513 TFLOPS1,670 TFLOPS
INT8 Tensor Core 1,979 TOPS1,513 TOPS1,670 TOPS
BF16 Tensor Core 990 TFLOPS756 TFLOPS835 TFLOPS
FP16 Tensor Core 990 TFLOPS756 TFLOPS835 TFLOPS
TF32 Tensor Core 494 TFLOPS378 TFLOPS417 TFLOPS
FP32 67 TFLOPS51 TFLOPS60 TFLOPS
FP64 34 TFLOPS26 TFLOPS30 TFLOPS
FP64 Tensor Core 67 TFLOPS51 TFLOPS60 TFLOPS

Performance figures assume no sparsity; in cases where only sparse performance figures are published by the manufacturer, these are halved to give approximate dense performance.

Resources

Detailed documentation from the manufacturer.

Similar GPUs

Other accelerators you might compare.