NVIDIA A100

Data center GPU · Ampere · SXM/PCIe, 80GB/40GB variants

Summary

NVIDIA's flagship Ampere-generation data center GPU. The A100's four variants share the same core Ampere architecture but differ in memory and interface, as well as memory bandwidth and power.

The A100 SXM 80GB is the highest-performance variant, and the most widely available. Together with the lower-memory A100 SXM 40GB, the A100's two SXM variants are the ones better suited to multi-GPU training, with fast interconnects between each of the GPUs.

The A100 PCIe 80GB and A100 PCIe 40GB can be a good fit for workloads requiring just one or two GPUs.

Launched
Q3 2020
VRAM
40 - 80 GB
Mem. bandwidth
1,555 - 2,039 GB/s
On-demand from
Compare prices from 21 providers →

Tech specs

Spec A100 80GB SXMA100 40GB SXMA100 80GB PCIeA100 40GB PCIe
VRAM 80 GB HBM2e40 GB HBM280 GB HBM2e40 GB HBM2
Memory bandwidth 2,039 GB/s1,555 GB/s1,935 GB/s1,555 GB/s
Interface SXM SXM PCIe PCIe
CUDA cores 6,912 6,912 6,912 6,912
Tensor cores 432 (Gen 3)432 (Gen 3)432 (Gen 3)432 (Gen 3)
TDP 400 W400 W300 W250 W
Supported data types
FP64FP32FP16BF16INT8INT4

Cloud rental prices

21 providers · available in 30 countries · lowest price per GPU, per hour

Prices updated

On-demand
$0.58 – $5.07 /GPU/h
Reserved
$1.36 – $2.35 /GPU/h
Spot
$0.63 – $1.21 /GPU/h
Compare all A100 80GB SXM cloud providers & configurations →
On-demand
$1.15 – $3.67 /GPU/h
Reserved
$1.17 – $2.31 /GPU/h
Spot
$0.32 – $0.45 /GPU/h
Compare all A100 40GB SXM cloud providers & configurations →
On-demand
$1.03 – $3.67 /GPU/h
Reserved
$0.57 – $3.42 /GPU/h
Spot
$0.68 – $1.08 /GPU/h
Compare all A100 80GB PCIe cloud providers & configurations →
On-demand
$0.89 – $1.99 /GPU/h
Reserved
$0.76 – $2.87 /GPU/h
Spot
$0.79 /GPU/h
Provider On-demand Reserved Spot
Cirrascale $2.30
Denvr $1.15
Jarvislabs $0.89 $0.76 $0.79
Lambda $1.99
Compare all A100 40GB PCIe cloud providers & configurations →

Models that fit in VRAM

Total VRAM From 16-bit inference 8-bit inference 4-bit inference
1× A100 80 GB $0.76/h GLM-4.7-FlashQwen3-Coder-30B-A3B-InstructNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 GLM-4.7-FlashQwen3-Coder-30B-A3B-InstructNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 gpt-oss-120bNVIDIA-Nemotron-3-Super-120B-A12B-BF16GLM-4.5-Air
2× A100 160 GB $1.52/h GLM-4.7-FlashQwen3-Coder-30B-A3B-InstructNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 gpt-oss-120bNVIDIA-Nemotron-3-Super-120B-A12B-BF16GLM-4.5-Air MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b
3× A100 240 GB $4.17/h GLM-4.5-AirQwen3-Coder-NextNVIDIA-Nemotron-3-Nano-30B-A3B-BF16 DeepSeek-V4-Flashgpt-oss-120bNVIDIA-Nemotron-3-Super-120B-A12B-BF16 MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b
4× A100 320 GB $3.04/h gpt-oss-120bNVIDIA-Nemotron-3-Super-120B-A12B-BF16GLM-4.5-Air MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b
5× A100 400 GB $6.95/h DeepSeek-V4-Flashgpt-oss-120bNVIDIA-Nemotron-3-Super-120B-A12B-BF16 MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b DeepSeek-V3.2MiniMax-M2.7gpt-oss-120b
6× A100 480 GB $8.34/h DeepSeek-V4-Flashgpt-oss-120bNVIDIA-Nemotron-3-Super-120B-A12B-BF16 MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b DeepSeek-V4-ProMiniMax-M2.7gpt-oss-120b
7× A100 560 GB $9.73/h MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b DeepSeek-V4-ProMiniMax-M2.7gpt-oss-120b
8× A100 640 GB $4.52/h MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b Kimi-K2-Instruct-0905DeepSeek-V4-ProMiniMax-M2.7
16× A100 1,280 GB $19.51/h MiniMax-M2.7DeepSeek-V4-Flashgpt-oss-120b Kimi-K2-Instruct-0905DeepSeek-V4-ProMiniMax-M2.7 Kimi-K2-Instruct-0905DeepSeek-V4-ProMiniMax-M2.7

Open-weights models that fit in VRAM. Estimated using 🤗 accelerate, plus approximation for up to 8K context.

Media

Peak theoretical performance

Spec A100 80GB SXMA100 40GB SXMA100 80GB PCIeA100 40GB PCIe
INT8 Tensor Core 624 TOPS624 TOPS624 TOPS624 TOPS
BF16 Tensor Core 312 TFLOPS312 TFLOPS312 TFLOPS312 TFLOPS
FP16 Tensor Core 312 TFLOPS312 TFLOPS312 TFLOPS312 TFLOPS
TF32 Tensor Core 156 TFLOPS156 TFLOPS156 TFLOPS156 TFLOPS
FP32 19.5 TFLOPS19.5 TFLOPS19.5 TFLOPS19.5 TFLOPS
FP64 9.7 TFLOPS9.7 TFLOPS9.7 TFLOPS9.7 TFLOPS
FP64 Tensor Core 19.5 TFLOPS19.5 TFLOPS19.5 TFLOPS19.5 TFLOPS

Performance figures assume no sparsity; in cases where only sparse performance figures are published by the manufacturer, these are halved to give approximate dense performance.

Resources

Detailed documentation from the manufacturer.

Similar GPUs

Other accelerators you might compare.