NVIDIA GH200
Data center GPU · Hopper · Grace
Launched
Q4 2024
VRAM
96 GB
Mem. bandwidth
4,000 GB/s
On-demand from
Tech specs
VRAM 96 GB HBM3
Memory bandwidth 4,000 GB/s
Interface PCIe
CUDA cores 16,896
Tensor cores 528 (Gen 4)
TDP 450 W
Supported data types
FP64FP32FP16BF16FP8INT8
Technical specifications per GPU; for whole system specs see manufacturer documentation.
Cloud rental prices
4 providers · available in 1 country · lowest price per GPU, per hour
Models that fit in VRAM
| Total VRAM | From | 16-bit inference | 8-bit inference | 4-bit inference | |
|---|---|---|---|---|---|
| 1× GH200 | 96 GB | $1.99/h | | | |
Open-weights models that fit in VRAM. Estimated using 🤗 accelerate, plus approximation for up to 8K context.
Media
Europe's First Exascale Supercomputer, JUPITER, Accelerates Climate Research, Neuroscience, Quantum Simulation Putting the NVIDIA GH200 Grace Hopper Superchip to good use: superior inference performance and economics for larger models NVIDIA GH200 Grace Hopper Superchips Now on Lambda and Available On-Demand Scaleway Uses NVIDIA GH200 Grace Hopper Superchip NVIDIA Grace Hopper Superchips Designed for Accelerated Generative AI Enter Full Production
Peak theoretical performance
FP8 Tensor Core 1,979 TFLOPS
INT8 Tensor Core 1,979 TOPS
BF16 Tensor Core 990 TFLOPS
FP16 Tensor Core 990 TFLOPS
TF32 Tensor Core 494 TFLOPS
FP32 67 TFLOPS
FP64 34 TFLOPS
FP64 Tensor Core 67 TFLOPS
Performance figures assume no sparsity; in cases where only sparse performance figures are published by the manufacturer, these are halved to give approximate dense performance.
Resources
Detailed documentation from the manufacturer.


