NVIDIA H200

Data center GPU · Hopper · SXM and NVL variants

Summary

An improved version of the H100, with additional memory and higher memory bandwidth, the H200 comes in two variants which share the same core Hopper architecture but differ in performance, form factor, interconnect, memory, and power.

The H200 SXM is the highest-performance variant, and the most widely available. It supports fast NVLink connectivity between GPUs, making it the best fit for multi-GPU training. The H200 NVL's lower power draw makes it easier to deploy and cool, but comes with performance trade-offs.

It is also available as part of the GH200 "Grace Hopper" system, which combines NVIDIA CPU and GPU chips.

Launched
Q3 2024
VRAM
141 GB
Mem. bandwidth
4,800 GB/s
On-demand from
Compare prices from 18 providers →

Tech specs

Spec H200 SXMH200 NVL
VRAM 141 GB HBM3e141 GB HBM3e
Memory bandwidth 4,800 GB/s4,800 GB/s
Interface SXM PCIe
CUDA cores 16,896 14,592
Tensor cores 528 (Gen 4)456 (Gen 4)
TDP 700 W600 W
Supported data types
FP64FP32FP16BF16FP8INT8

Cloud rental prices

18 providers · available in 26 countries · lowest price per GPU, per hour

Prices updated

On-demand
$2.97 – $10.60 /GPU/h
Reserved
$2.53 – $7.31 /GPU/h
Spot
$1.40 – $5.36 /GPU/h
Compare all H200 SXM cloud providers & configurations →
On-demand
$3.44 – $3.82 /GPU/h
Reserved
$3.43 – $4.54 /GPU/h
Provider On-demand Reserved Spot
Cirrascale $3.43
DigitalOcean $3.44
Massed Compute $3.62
Compare all H200 NVL cloud providers & configurations →

Models that fit in VRAM

Open-weights models that fit in VRAM. Estimated using 🤗 accelerate, plus approximation for up to 8K context.

Media

Peak theoretical performance

Spec H200 SXMH200 NVL
FP8 Tensor Core 1,979 TFLOPS1,670 TFLOPS
INT8 Tensor Core 1,979 TOPS1,670 TOPS
BF16 Tensor Core 990 TFLOPS835 TFLOPS
FP16 Tensor Core 990 TFLOPS835 TFLOPS
TF32 Tensor Core 494 TFLOPS417 TFLOPS
FP32 67 TFLOPS60 TFLOPS
FP64 34 TFLOPS30 TFLOPS
FP64 Tensor Core 67 TFLOPS60 TFLOPS

Performance figures assume no sparsity; in cases where only sparse performance figures are published by the manufacturer, these are halved to give approximate dense performance.

Resources

Detailed documentation from the manufacturer.

Similar GPUs

Other accelerators you might compare.