Introduction: The DGX H100 in Context

The NVIDIA DGX H100 is the flagship enterprise AI server, delivering 32 petaFLOPS of FP8 AI performance from eight H100 SXM5 GPUs interconnected via NVLink 4.0 at 900 GB/s bidirectional bandwidth. Introduced in 2022 and shipping broadly through 2023, the DGX H100 is built around the Hopper architecture H100 SXM5 GPU — the same silicon that powers the world's most capable AI training clusters.

For enterprise AI teams, the DGX H100 occupies a specific position in the infrastructure hierarchy: it is the highest-performance single-node AI server available, and it is the building block for DGX SuperPOD clusters. Understanding its capabilities, constraints, and operational requirements is essential for any organization planning large-scale AI training or high-throughput inference infrastructure.

32 PF

FP8 AI performance per DGX H100 system

640 GB

Total HBM2e GPU memory across 8 GPUs

900 GB/s

NVLink 4.0 bidirectional bandwidth

10.2 kW

System TDP — requires liquid cooling

Technical Specifications

The DGX H100 is a 10U chassis housing eight H100 SXM5 GPUs, dual Intel Xeon Platinum 8480C CPUs, 2TB of system DRAM, and eight NVMe SSDs. The SXM5 form factor is critical — it enables the high-bandwidth NVLink connections that are not possible with PCIe-based GPU installations.

SXM5 vs PCIe Form Factor

The H100 SXM5 GPU used in the DGX H100 is not the same product as the H100 PCIe card. SXM5 GPUs are soldered directly to a high-bandwidth substrate and connect to NVSwitches via NVLink. This enables 900 GB/s GPU-to-GPU bandwidth — compared to 64 GB/s for PCIe 5.0 x16. The SXM5 form factor also requires liquid cooling; there is no air-cooled SXM5 option.
GPU8x NVIDIA H100 SXM5 80GB
GPU Memory640 GB HBM2e total (80 GB per GPU)
GPU Memory BW3.35 TB/s aggregate
AI Performance32 petaFLOPS FP8
CPU2x Intel Xeon Platinum 8480C (56 cores each)
System Memory2 TB DDR5-4800
Storage8x 3.84 TB NVMe SSD (30.72 TB total)
Network8x 400GbE/InfiniBand HDR + 2x 10GbE mgmt
Form Factor10U, 19-inch rack
Weight132 kg (291 lbs)
Power10.2 kW typical, 11.5 kW max
CoolingLiquid cooling required (rear-door or direct liquid)

NVLink 4.0 Architecture and NVSwitch Topology

The DGX H100's performance advantage over PCIe GPU servers comes entirely from its NVLink 4.0 fabric. Four NVSwitch 3.0 chips create a fully connected all-to-all topology between all eight GPUs, enabling any GPU to communicate with any other GPU at full 900 GB/s bidirectional bandwidth simultaneously.

DGX H100 Internal Architecture

Network and Storage

External connectivity

8x 400GbE/IB2x 10GbE Mgmt8x NVMe SSDBMC/IPMI

CPU and Memory

Dual Intel Xeon + 2TB DDR5

CPU 0 (56c)CPU 1 (56c)System DRAM 2TBPCIe 5.0 Fabric

NVLink Fabric

4x NVSwitch 3.0 — 900 GB/s bidirectional

NVSwitch 0NVSwitch 1NVSwitch 2NVSwitch 3

GPU Compute Layer

8x H100 SXM5 GPUs

GPU 0GPU 1GPU 2GPU 3GPU 4GPU 5GPU 6GPU 7
Stack layers — top to bottom: highest to lowest abstraction

NVLink 4.0 provides 18 links per GPU, each operating at 50 Gb/s bidirectional. With four NVSwitches providing full any-to-any connectivity, the aggregate NVLink bandwidth within the DGX H100 is 7.2 TB/s — sufficient to sustain tensor parallel operations across all eight GPUs without communication becoming the bottleneck.

DGX Platform Comparison

The DGX product line spans multiple GPU generations. Understanding the differences between DGX A100, DGX H100, DGX H200, and the HGX H100 platform helps organizations make the right procurement decision for their workload and budget.

DGX H100 vs DGX A100 vs DGX H200 vs HGX H100

DimensionDGX A100DGX H100DGX H200HGX H100
GPUs8x A100 SXM48x H100 SXM58x H200 SXM58x H100 SXM5
GPU Memory640 GB HBM2e640 GB HBM2e1,128 GB HBM3e640 GB HBM2e
GPU Interconnect BW600 GB/s900 GB/s900 GB/s900 GB/s
AI Performance (FP8)5 PF32 PF32 PF32 PF
Power6.5 kW10.2 kW10.2 kW10.2 kW
CoolingAir or liquidLiquid requiredLiquid requiredLiquid required
List Price~$200K~$300–400K~$350–450K~$250–350K
Best ForExisting deploymentsTraining + inferenceLarge model inferenceOEM integration

HGX vs DGX: What Is the Difference?

The HGX H100 is the GPU board assembly sold to OEM server vendors like Dell, HPE, and Supermicro. The DGX H100 is NVIDIA's complete, vertically integrated system including CPUs, memory, storage, networking, and software stack. HGX-based servers are typically 10–20% less expensive but require more integration work and may have different software support timelines.

Deployment Requirements

Deploying a DGX H100 requires significant facility preparation. The system's 10.2 kW typical power draw and mandatory liquid cooling requirement mean that most standard data center environments need upgrades before a DGX H100 can be installed. Planning these facility requirements 6–12 months in advance is essential given lead times for electrical and cooling infrastructure.

  1. 1

    Power Circuit Requirements

    The DGX H100 requires a dedicated 200–240V, 60A single-phase or 30A 3-phase circuit. A standard 20A circuit is insufficient. Most deployments use a 60A 3-phase PDU with C19 outlets. Verify that your facility's electrical panel has capacity before ordering.

  2. 2

    Liquid Cooling Installation

    The DGX H100 uses a rear-door heat exchanger (RDHx) or direct liquid cooling (DLC) manifold. RDHx is the most common approach — it attaches to the rear of the rack and connects to facility chilled water. DLC provides higher efficiency but requires more complex plumbing. Both require chilled water supply at 18–24°C.

  3. 3

    Rack and Floor Loading

    At 132 kg (291 lbs), the DGX H100 requires a rack rated for at least 1,500 kg total load and a raised floor or structural floor rated for concentrated point loads. Verify floor loading capacity with your facilities team before installation.

  4. 4

    Network Infrastructure

    Each DGX H100 has 8x 400GbE/InfiniBand HDR ports for GPU fabric connectivity and 2x 10GbE ports for management. You will need 400GbE or InfiniBand HDR switches and appropriate cables. Standard 1GbE switches are insufficient for GPU fabric.

  5. 5

    Software Stack

    NVIDIA provides the DGX OS (Ubuntu-based) and the DGX software stack including drivers, CUDA, cuDNN, NCCL, and the NGC container catalog. Plan for a 4–8 hour initial software configuration process after physical installation.

Networking and Cluster Connectivity

The DGX H100's eight 400GbE/InfiniBand ports are designed for direct connection to InfiniBand NDR or 400GbE switches, enabling DGX H100 systems to be interconnected into DGX SuperPOD clusters. Each of the eight ports connects to one of the eight H100 GPUs, providing direct GPU-to-network connectivity without CPU involvement in data transfer.

Rail-Optimized Networking Required at Scale

When connecting multiple DGX H100 systems, use rail-optimized networking — each GPU's network port connects to a dedicated leaf switch, and all leaf switches connect to spine switches. This topology ensures that all-to-all GPU communication achieves full bisection bandwidth. Non-rail-optimized topologies create bottlenecks that severely degrade training performance above 4 nodes.
400 Gb/s

Per-port bandwidth (IB NDR or 400GbE)

8 ports

GPU fabric ports per DGX H100

3.2 Tb/s

Total GPU fabric bandwidth per node

600 ns

InfiniBand NDR latency

Supported Workloads and Performance Benchmarks

The DGX H100 is optimized for large-scale AI training and high-throughput inference. Its 640 GB of HBM2e memory enables single-node training and inference for models up to approximately 300B parameters (with quantization), while its NVLink fabric makes tensor parallelism across all eight GPUs practical.

LLM Training (7B–70B parameters)

Single DGX H100 can train LLaMA-70B with 8-way tensor parallelism

Optimal for models fitting in 640GB GPU memory

LLM Inference (7B–70B parameters)

Up to 30,000 tokens/sec throughput for LLaMA-70B at FP16

Tensor parallelism across 8 GPUs for low latency

Computer Vision Training

8x faster than A100 for ResNet-50 training

FP8 precision provides 2x speedup over FP16

Recommendation Systems

High memory bandwidth suits embedding table lookups

HBM2e bandwidth critical for sparse access patterns

Scientific Simulation

FP64 performance: 60 TFLOPS per GPU

Suitable for molecular dynamics, climate modeling

ROI and Business Case

The DGX H100 business case depends on utilization rate and the revenue or cost savings generated per GPU-hour. At cloud GPU spot prices of $3–$8/GPU-hour, a fully utilized DGX H100 can generate $210K–$560K per year in avoided cloud costs — providing payback in 12–24 months at typical enterprise utilization rates.

DGX H100 ROI Calculator

Estimate annual revenue, payback period, and 3-year ROI for DGX H100 deployments based on utilization and GPU-hour pricing.

2 units
120
350,000 $
300,000400,000
75 %
5095
6 $/hr
315

Estimated results

$0.70M

Total CapEx

$631K

Annual GPU Revenue

15 months

Payback Period

146%

3-Year ROI

Operations and Management

DGX H100 systems are managed through NVIDIA's DGX software stack, which includes DCGM for health monitoring, the DGX OS for system management, and NGC for container-based workload deployment. Operational teams should plan for regular firmware updates, ECC memory monitoring, and liquid cooling maintenance.

DCGM Is Your Primary Monitoring Tool

NVIDIA DCGM provides GPU health metrics, XID error monitoring, ECC memory tracking, and performance telemetry. Integrate DCGM with your existing Prometheus/Grafana stack using the DCGM Exporter. Set alerts for GPU temperature above 83°C, ECC double-bit errors, and XID errors 79 and 94.
83°C

Maximum GPU junction temperature

5 yr

Typical DGX H100 operational lifespan

DCGM

Primary GPU health monitoring tool

4 hr

Typical firmware update maintenance window

Frequently Asked Questions

What is the difference between DGX H100 and HGX H100?

The DGX H100 is NVIDIA's complete, vertically integrated AI server — it includes CPUs, system memory, storage, networking, and the full DGX software stack in addition to the GPU assembly. The HGX H100 is just the GPU board assembly sold to OEM server vendors. DGX systems typically have better software support and faster time-to-deployment; HGX-based systems offer more configuration flexibility and are often 10–20% less expensive.

Does DGX H100 require liquid cooling?

Yes, liquid cooling is mandatory for the DGX H100. The H100 SXM5 GPU has a TDP of 700W per GPU, and with eight GPUs plus CPUs and other components, the system dissipates over 10 kW. Air cooling cannot remove this heat density in a 10U chassis. NVIDIA supports rear-door heat exchangers and direct liquid cooling configurations. Facilities must provide chilled water at 18–24°C before DGX H100 installation.

How many DGX H100s do you need for LLM training?

It depends on the model size and training timeline. A single DGX H100 (640 GB GPU memory) can train models up to approximately 70B parameters. For 175B parameter models, you need at least 2 DGX H100 systems. For 405B parameter models, plan for 4–8 systems. For training frontier models (1T+ parameters), you need 64+ DGX H100 systems in a SuperPOD configuration.

What is the list price of DGX H100?

NVIDIA's list price for the DGX H100 is approximately $300,000–$400,000 depending on configuration and region. Enterprise customers purchasing through NVIDIA-authorized partners typically receive 5–15% discounts on multi-unit orders. Total cost of ownership including facility preparation, networking, and 3-year support contracts typically adds 30–50% to the hardware list price.