Ampere
A mature baseline for CUDA 11-era deployments, TF32, and first-generation MIG. Useful when an established software image matters more than the newest precision formats.
GPU Compare
Select up to three GPUs. Specifications are reference values from NVIDIA; final configuration depends on the exact platform and cooling design.
3 GPUs selected
| Specification | HopperH200 SXM | HopperH200 NVL | HopperH100 SXM | HopperH100 NVL | BlackwellRTX PRO 6000 | Ada LovelaceL40S | AmpereA100 80GB SXM |
|---|---|---|---|---|---|---|---|
| ArchitectureGeneration affects supported precision, tensor features, and software capabilities. | Hopper | Hopper | Hopper | Hopper | Blackwell | Ada Lovelace | Ampere |
| GPU memoryCapacity determines whether the model, working set, and runtime overhead fit on each GPU. | 141 GB HBM3e | 141 GB HBM3e | 80 GB HBM3 | 94 GB HBM3 | 96 GB GDDR7 ECC | 48 GB GDDR6 ECC | 80 GB HBM2e |
| Memory bandwidthBandwidth can govern throughput when a workload repeatedly moves large tensors or datasets. | 4.8 TB/s | 4.8 TB/s | 3.35 TB/s | 3.9 TB/s | 1.597 TB/s | 864 GB/s | 2.039 TB/s |
| Maximum powerPower affects chassis choice, rack density, cooling, and facility planning. | Up to 700 W | Up to 600 W | Up to 700 W | 350–400 W | Up to 600 W | 350 W | 400 W |
| Form factorSXM and PCIe accelerators require different platforms and expansion strategies. | SXM | PCIe · dual slot | SXM | PCIe · dual slot | PCIe · air or liquid cooled | PCIe · dual slot | SXM |
| Host / GPU fabricInterconnect choice becomes increasingly important for multi-GPU and multi-node workloads. | NVLink 900 GB/s · PCIe Gen5 | 2- or 4-way NVLink bridge · PCIe Gen5 | NVLink 900 GB/s · PCIe Gen5 | NVLink 600 GB/s · PCIe Gen5 | PCIe Gen5 | PCIe Gen4 · no NVLink | NVLink 600 GB/s · PCIe Gen4 |
| MIG supportHardware partitioning can improve isolation and utilization for shared environments. | Up to 7 × 18 GB | Up to 7 × 16.5 GB | Up to 7 instances | Up to 7 instances | Up to 4 instances | Not supported | Up to 7 × 10 GB |
| Good starting point forA workload fit is a starting point, not a substitute for application-level validation. | Memory-bound LLM training, inference, and HPC in HGX systems | High-memory AI inference and HPC in MGX-class PCIe platforms | Dense AI training and HPC in HGX-class systems | High-memory LLM inference in PCIe platforms | Enterprise AI, rendering, simulation, and visual compute | Inference, rendering, video, and virtual workstation workloads | Established AI and HPC environments with mature Ampere software stacks |
| Official sourceOpen the manufacturer page for product notes and full specifications. | NVIDIA product page ↗ | NVIDIA product page ↗ | NVIDIA product page ↗ | NVIDIA product page ↗ | NVIDIA product page ↗ | NVIDIA product page ↗ | NVIDIA product page ↗ |
Architecture and framework fit
PyTorch, JAX, TensorFlow, and inference runtimes can target multiple GPU generations, but the exact driver, CUDA toolkit, framework build, kernels, precision, and container image still need validation.
A mature baseline for CUDA 11-era deployments, TF32, and first-generation MIG. Useful when an established software image matters more than the newest precision formats.
Strong visual-compute, ray-tracing, and media capabilities alongside AI inference. Products such as L40S target a different system profile from HGX training accelerators.
Adds Transformer Engine with FP8, fourth-generation NVLink, second-generation MIG, confidential computing, and DPX instructions for selected HPC algorithms.
Adds newer low-precision paths including FP4-class capabilities on supported products. Confirm driver, CUDA, framework, and kernel support before treating peak capability as delivered application performance.
Check supported CUDA builds, distributed backends, precision paths, and custom operators.
Kernel coverage, quantization support, batching behavior, and architecture maturity can change delivered throughput.
A newer GPU may require a newer driver or base image even when application code remains unchanged.
Selected a direction?
GPU choice affects CPU balance, memory, storage, networking, chassis, cooling, and power. Send the selected models for a platform review.
Ask about selected GPUs