Introduction: The DGX H100 in Context
The NVIDIA DGX H100 is the flagship enterprise AI server, delivering 32 petaFLOPS of FP8 AI performance from eight H100 SXM5 GPUs interconnected via NVLink 4.0 at 900 GB/s bidirectional bandwidth. Introduced in 2022 and shipping broadly through 2023, the DGX H100 is built around the Hopper architecture H100 SXM5 GPU — the same silicon that powers the world's most capable AI training clusters.
For enterprise AI teams, the DGX H100 occupies a specific position in the infrastructure hierarchy: it is the highest-performance single-node AI server available, and it is the building block for DGX SuperPOD clusters. Understanding its capabilities, constraints, and operational requirements is essential for any organization planning large-scale AI training or high-throughput inference infrastructure.
FP8 AI performance per DGX H100 system
Total HBM2e GPU memory across 8 GPUs
NVLink 4.0 bidirectional bandwidth
System TDP — requires liquid cooling
Technical Specifications
The DGX H100 is a 10U chassis housing eight H100 SXM5 GPUs, dual Intel Xeon Platinum 8480C CPUs, 2TB of system DRAM, and eight NVMe SSDs. The SXM5 form factor is critical — it enables the high-bandwidth NVLink connections that are not possible with PCIe-based GPU installations.
SXM5 vs PCIe Form Factor
NVLink 4.0 Architecture and NVSwitch Topology
The DGX H100's performance advantage over PCIe GPU servers comes entirely from its NVLink 4.0 fabric. Four NVSwitch 3.0 chips create a fully connected all-to-all topology between all eight GPUs, enabling any GPU to communicate with any other GPU at full 900 GB/s bidirectional bandwidth simultaneously.
DGX H100 Internal Architecture
Network and Storage
External connectivity
CPU and Memory
Dual Intel Xeon + 2TB DDR5
NVLink Fabric
4x NVSwitch 3.0 — 900 GB/s bidirectional
GPU Compute Layer
8x H100 SXM5 GPUs
NVLink 4.0 provides 18 links per GPU, each operating at 50 Gb/s bidirectional. With four NVSwitches providing full any-to-any connectivity, the aggregate NVLink bandwidth within the DGX H100 is 7.2 TB/s — sufficient to sustain tensor parallel operations across all eight GPUs without communication becoming the bottleneck.
DGX Platform Comparison
The DGX product line spans multiple GPU generations. Understanding the differences between DGX A100, DGX H100, DGX H200, and the HGX H100 platform helps organizations make the right procurement decision for their workload and budget.
DGX H100 vs DGX A100 vs DGX H200 vs HGX H100
| Dimension | DGX A100 | DGX H100 | DGX H200 | HGX H100 |
|---|---|---|---|---|
| GPUs | 8x A100 SXM4 | 8x H100 SXM5 | 8x H200 SXM5 | 8x H100 SXM5 |
| GPU Memory | 640 GB HBM2e | 640 GB HBM2e | 1,128 GB HBM3e | 640 GB HBM2e |
| GPU Interconnect BW | 600 GB/s | 900 GB/s | 900 GB/s | 900 GB/s |
| AI Performance (FP8) | 5 PF | 32 PF | 32 PF | 32 PF |
| Power | 6.5 kW | 10.2 kW | 10.2 kW | 10.2 kW |
| Cooling | Air or liquid | Liquid required | Liquid required | Liquid required |
| List Price | ~$200K | ~$300–400K | ~$350–450K | ~$250–350K |
| Best For | Existing deployments | Training + inference | Large model inference | OEM integration |
HGX vs DGX: What Is the Difference?
Deployment Requirements
Deploying a DGX H100 requires significant facility preparation. The system's 10.2 kW typical power draw and mandatory liquid cooling requirement mean that most standard data center environments need upgrades before a DGX H100 can be installed. Planning these facility requirements 6–12 months in advance is essential given lead times for electrical and cooling infrastructure.
- 1
Power Circuit Requirements
The DGX H100 requires a dedicated 200–240V, 60A single-phase or 30A 3-phase circuit. A standard 20A circuit is insufficient. Most deployments use a 60A 3-phase PDU with C19 outlets. Verify that your facility's electrical panel has capacity before ordering.
- 2
Liquid Cooling Installation
The DGX H100 uses a rear-door heat exchanger (RDHx) or direct liquid cooling (DLC) manifold. RDHx is the most common approach — it attaches to the rear of the rack and connects to facility chilled water. DLC provides higher efficiency but requires more complex plumbing. Both require chilled water supply at 18–24°C.
- 3
Rack and Floor Loading
At 132 kg (291 lbs), the DGX H100 requires a rack rated for at least 1,500 kg total load and a raised floor or structural floor rated for concentrated point loads. Verify floor loading capacity with your facilities team before installation.
- 4
Network Infrastructure
Each DGX H100 has 8x 400GbE/InfiniBand HDR ports for GPU fabric connectivity and 2x 10GbE ports for management. You will need 400GbE or InfiniBand HDR switches and appropriate cables. Standard 1GbE switches are insufficient for GPU fabric.
- 5
Software Stack
NVIDIA provides the DGX OS (Ubuntu-based) and the DGX software stack including drivers, CUDA, cuDNN, NCCL, and the NGC container catalog. Plan for a 4–8 hour initial software configuration process after physical installation.
Networking and Cluster Connectivity
The DGX H100's eight 400GbE/InfiniBand ports are designed for direct connection to InfiniBand NDR or 400GbE switches, enabling DGX H100 systems to be interconnected into DGX SuperPOD clusters. Each of the eight ports connects to one of the eight H100 GPUs, providing direct GPU-to-network connectivity without CPU involvement in data transfer.
Rail-Optimized Networking Required at Scale
Per-port bandwidth (IB NDR or 400GbE)
GPU fabric ports per DGX H100
Total GPU fabric bandwidth per node
InfiniBand NDR latency
Supported Workloads and Performance Benchmarks
The DGX H100 is optimized for large-scale AI training and high-throughput inference. Its 640 GB of HBM2e memory enables single-node training and inference for models up to approximately 300B parameters (with quantization), while its NVLink fabric makes tensor parallelism across all eight GPUs practical.
LLM Training (7B–70B parameters)
Single DGX H100 can train LLaMA-70B with 8-way tensor parallelism
Optimal for models fitting in 640GB GPU memory
LLM Inference (7B–70B parameters)
Up to 30,000 tokens/sec throughput for LLaMA-70B at FP16
Tensor parallelism across 8 GPUs for low latency
Computer Vision Training
8x faster than A100 for ResNet-50 training
FP8 precision provides 2x speedup over FP16
Recommendation Systems
High memory bandwidth suits embedding table lookups
HBM2e bandwidth critical for sparse access patterns
Scientific Simulation
FP64 performance: 60 TFLOPS per GPU
Suitable for molecular dynamics, climate modeling
ROI and Business Case
The DGX H100 business case depends on utilization rate and the revenue or cost savings generated per GPU-hour. At cloud GPU spot prices of $3–$8/GPU-hour, a fully utilized DGX H100 can generate $210K–$560K per year in avoided cloud costs — providing payback in 12–24 months at typical enterprise utilization rates.
DGX H100 ROI Calculator
Estimate annual revenue, payback period, and 3-year ROI for DGX H100 deployments based on utilization and GPU-hour pricing.
Estimated results
Total CapEx
Annual GPU Revenue
Payback Period
3-Year ROI
Operations and Management
DGX H100 systems are managed through NVIDIA's DGX software stack, which includes DCGM for health monitoring, the DGX OS for system management, and NGC for container-based workload deployment. Operational teams should plan for regular firmware updates, ECC memory monitoring, and liquid cooling maintenance.
DCGM Is Your Primary Monitoring Tool
Maximum GPU junction temperature
Typical DGX H100 operational lifespan
Primary GPU health monitoring tool
Typical firmware update maintenance window
Frequently Asked Questions
What is the difference between DGX H100 and HGX H100?
The DGX H100 is NVIDIA's complete, vertically integrated AI server — it includes CPUs, system memory, storage, networking, and the full DGX software stack in addition to the GPU assembly. The HGX H100 is just the GPU board assembly sold to OEM server vendors. DGX systems typically have better software support and faster time-to-deployment; HGX-based systems offer more configuration flexibility and are often 10–20% less expensive.
Does DGX H100 require liquid cooling?
Yes, liquid cooling is mandatory for the DGX H100. The H100 SXM5 GPU has a TDP of 700W per GPU, and with eight GPUs plus CPUs and other components, the system dissipates over 10 kW. Air cooling cannot remove this heat density in a 10U chassis. NVIDIA supports rear-door heat exchangers and direct liquid cooling configurations. Facilities must provide chilled water at 18–24°C before DGX H100 installation.
How many DGX H100s do you need for LLM training?
It depends on the model size and training timeline. A single DGX H100 (640 GB GPU memory) can train models up to approximately 70B parameters. For 175B parameter models, you need at least 2 DGX H100 systems. For 405B parameter models, plan for 4–8 systems. For training frontier models (1T+ parameters), you need 64+ DGX H100 systems in a SuperPOD configuration.
What is the list price of DGX H100?
NVIDIA's list price for the DGX H100 is approximately $300,000–$400,000 depending on configuration and region. Enterprise customers purchasing through NVIDIA-authorized partners typically receive 5–15% discounts on multi-unit orders. Total cost of ownership including facility preparation, networking, and 3-year support contracts typically adds 30–50% to the hardware list price.