System Architecture
An AI infrastructure system has five interdependent layers: compute (GPU nodes), interconnect (network fabric), storage (parallel file systems), power (PDUs, UPS, generators), and cooling (liquid or air). Each layer must be engineered to match the requirements of the others, a mismatch at any layer creates a bottleneck that limits the entire system.
The system design principle
Compute Layer
The compute layer consists of GPU servers, each containing 4–8 GPUs connected via NVLink (for intra-node GPU-to-GPU communication) and PCIe (for CPU-to-GPU and NIC attachment). Current generation AI servers use NVIDIA H100 or H200 GPUs in DGX, HGX, or OEM configurations.
Current Generation AI Server Platforms
| Platform | GPUs | GPU Memory | NVLink | Power Draw | Form Factor |
|---|---|---|---|---|---|
| NVIDIA DGX H100 | 8× H100 SXM5 | 640 GB HBM3 | NVLink 4.0 (900 GB/s) | 10.2 kW | 6U rack |
| NVIDIA DGX H200 | 8× H200 SXM5 | 1.1 TB HBM3e | NVLink 4.0 (900 GB/s) | 10.2 kW | 6U rack |
| NVIDIA HGX H100 | 8× H100 SXM5 | 640 GB HBM3 | NVLink 4.0 | 10.2 kW | OEM chassis |
| OEM 4× GPU Server | 4× H100/H200 PCIe | 320 GB HBM3 | NVLink 3.0 (600 GB/s) | 5–6 kW | 2U rack |
Network Fabric
The network fabric connects GPU nodes for distributed training and connects the cluster to storage and management networks. The fabric must provide sufficient bandwidth to prevent GPU idle time during all-reduce operations, the collective communication pattern used in distributed training.
AI Network Fabric Options
| Fabric | Bandwidth | Latency | Scale | Cost | Best For |
|---|---|---|---|---|---|
| InfiniBand NDR | 400 Gb/s per port | ~500 ns | 1,000+ GPUs | High | Large training clusters, frontier models |
| InfiniBand HDR | 200 Gb/s per port | ~600 ns | 500+ GPUs | Medium-high | Mid-scale training, HPC |
| Ethernet 400 GbE + RoCEv2 | 400 Gb/s per port | 1–3 µs | Unlimited | Medium | Large inference, cost-sensitive training |
| Ethernet 100 GbE + RoCEv2 | 100 Gb/s per port | 1–5 µs | Unlimited | Low-medium | Small clusters, inference, dev/test |
Lossless fabric requirement
Storage Architecture
AI training requires storage that can feed data to GPUs faster than they can consume it. A single DGX H100 system can consume data at 200+ GB/s during training. A cluster of 64 DGX systems requires storage capable of delivering 12+ TB/s of aggregate throughput, far beyond what traditional enterprise storage can provide.
Hot tier, Parallel file system
Technologies: GPFS, Lustre, WEKA, VAST, DDN
Active training datasets, checkpoints, model weights. Must deliver hundreds of GB/s of sequential throughput.
Warm tier, High-capacity NVMe
Technologies: All-NVMe arrays, NVMe-oF
Preprocessed datasets, recent checkpoints, inference model storage.
Cold tier, Object storage
Technologies: S3-compatible, Ceph, MinIO
Raw datasets, archived models, long-term checkpoint storage.
Power Engineering
Power engineering for AI infrastructure requires sizing for actual GPU utilization under training load, not nameplate ratings. A DGX H100 system draws 10.2 kW at full load. A rack of 4 DGX systems draws 40+ kW: requiring 60A 3-phase circuits, high-density PDUs, and UPS systems sized for the full cluster.
Power capacity is the most common constraint
Cooling Systems
Air cooling cannot remove heat from racks drawing 20+ kW at the densities AI clusters require. The three liquid cooling approaches for AI infrastructure are rear-door heat exchangers (RDHx), direct-to-chip liquid cooling (DLC), and full immersion cooling.
Cooling Technology Comparison
| Technology | Max Rack Density | Cooling Efficiency | Infrastructure Change | Cost |
|---|---|---|---|---|
| Air cooling | Up to 15 kW | PUE 1.4–2.0 | None | Low |
| Rear-door heat exchanger | Up to 30 kW | PUE 1.2–1.5 | Chilled water to rack | Medium |
| Direct-to-chip liquid | Up to 60 kW | PUE 1.1–1.3 | Liquid loop to each server | High |
| Full immersion | Up to 100+ kW | PUE 1.02–1.1 | Complete facility redesign | Very high |
Design Patterns
Pod-based scaling
Design the cluster as a set of identical pods: each pod containing a fixed number of GPU nodes, network switches, and storage nodes. Scale by adding pods, not individual components.
Separate training and inference fabrics
Training requires high-bandwidth, low-latency InfiniBand. Inference requires high-throughput, low-latency Ethernet. Separate fabrics prevent training traffic from affecting inference latency.
Out-of-band management network
A dedicated management network (BMC, IPMI, iDRAC) separate from the data fabric allows infrastructure management without affecting training workloads.
Checkpoint storage co-location
Place checkpoint storage (NVMe) physically close to compute nodes to minimize checkpoint write latency, a critical factor for training job recovery time.