NVIDIA H200
Data center GPU · Hopper · SXM and NVL variants
An improved version of the H100, with additional memory and higher memory bandwidth, the H200 comes in two variants which share the same core Hopper architecture but differ in performance, form factor, interconnect, memory, and power.
The H200 SXM is the highest-performance variant, and the most widely available. It supports fast NVLink connectivity between GPUs, making it the best fit for multi-GPU training. The H200 NVL's lower power draw makes it easier to deploy and cool, but comes with performance trade-offs.
It is also available as part of the GH200 "Grace Hopper" system, which combines NVIDIA CPU and GPU chips.
Tech specs
| Spec | H200 SXM | H200 NVL |
|---|---|---|
| VRAM | 141 GB HBM3e | 141 GB HBM3e |
| Memory bandwidth | 4,800 GB/s | 4,800 GB/s |
| Interface | SXM | PCIe |
| CUDA cores | 16,896 | 14,592 |
| Tensor cores | 528 (Gen 4) | 456 (Gen 4) |
| TDP | 700 W | 600 W |
Cloud rental prices
18 providers · available in 26 countries · lowest price per GPU, per hour
| Provider | On-demand | Reserved | Spot |
|---|---|---|---|
AceCloud | — | $3.61 | — |
| | $7.91 | — | — |
| | $10.60 | $4.66 | $2.12 |
CoreWeave | $6.31 | — | $2.58 |
Crusoe | $4.29 | — | — |
| | $10.60 | $4.65 | $5.36 |
Hyperstack | $3.99 | $2.79 | — |
| | $4.50 | — | $2.45 |
| | $10.00 | — | — |
Runpod | $4.39 | — | — |
Seeweb | $2.97 | $2.53 | — |
| | $3.78 | — | — |
| | $5.99 | — | — |
Vast.ai | $4.08 | — | $3.48 |
| | $4.00 | $3.68 | $1.40 |
| Provider | On-demand | Reserved | Spot |
|---|---|---|---|
Cirrascale | — | $3.43 | — |
DigitalOcean | $3.44 | — | — |
Massed Compute | $3.62 | — | — |
Models that fit in VRAM
Open-weights models that fit in VRAM. Estimated using 🤗 accelerate, plus approximation for up to 8K context.
Media
Peak theoretical performance
| Spec | H200 SXM | H200 NVL |
|---|---|---|
| FP8 Tensor Core | 1,979 TFLOPS | 1,670 TFLOPS |
| INT8 Tensor Core | 1,979 TOPS | 1,670 TOPS |
| BF16 Tensor Core | 990 TFLOPS | 835 TFLOPS |
| FP16 Tensor Core | 990 TFLOPS | 835 TFLOPS |
| TF32 Tensor Core | 494 TFLOPS | 417 TFLOPS |
| FP32 | 67 TFLOPS | 60 TFLOPS |
| FP64 | 34 TFLOPS | 30 TFLOPS |
| FP64 Tensor Core | 67 TFLOPS | 60 TFLOPS |
Performance figures assume no sparsity; in cases where only sparse performance figures are published by the manufacturer, these are halved to give approximate dense performance.
Resources
Detailed documentation from the manufacturer.
Similar GPUs
Other accelerators you might compare.









