Skip to main content
DCS Global

AI Infrastructure Technical Overview

Technical IT leader / Technical evaluator 20 min

AI Infrastructure: Architecture, Components, and Engineering Decisions

The technical architecture of AI infrastructure: GPU cluster design, network fabric, storage systems, power engineering, and the system-level decisions that determine whether the infrastructure performs as designed.

Executive Summary

AI infrastructure is a tightly coupled system where every component affects every other. GPU utilization is limited by network bandwidth. Network performance is limited by switch configuration. Storage throughput is limited by file system design. Power density is limited by cooling capacity. Engineering AI infrastructure requires understanding these dependencies and designing the system as a whole, not selecting components independently.

Key Takeaways

  • GPU cluster performance is determined by the weakest link in the system: compute, network, storage, or power.
  • InfiniBand NDR (400 Gb/s) is the current standard for large AI training clusters; RoCEv2 is viable for smaller clusters and inference.
  • Parallel file systems (GPFS, Lustre, WEKA, VAST) are required for training workloads, NAS and SAN are insufficient.
  • Power density planning must account for GPU TDP at full utilization, not nameplate ratings, actual draw is typically 80–95% of TDP under training load.
  • Liquid cooling is required for racks above 20 kW: rear-door heat exchangers for moderate densities, direct-to-chip for high densities.

System Architecture

An AI infrastructure system has five interdependent layers: compute (GPU nodes), interconnect (network fabric), storage (parallel file systems), power (PDUs, UPS, generators), and cooling (liquid or air). Each layer must be engineered to match the requirements of the others, a mismatch at any layer creates a bottleneck that limits the entire system.

The system design principle

Design AI infrastructure as a system, not a collection of components. The performance of the system is determined by the interaction between layers, not by the peak performance of any individual component.

Compute Layer

The compute layer consists of GPU servers, each containing 4–8 GPUs connected via NVLink (for intra-node GPU-to-GPU communication) and PCIe (for CPU-to-GPU and NIC attachment). Current generation AI servers use NVIDIA H100 or H200 GPUs in DGX, HGX, or OEM configurations.

Current Generation AI Server Platforms

PlatformGPUsGPU MemoryNVLinkPower DrawForm Factor
NVIDIA DGX H1008× H100 SXM5640 GB HBM3NVLink 4.0 (900 GB/s)10.2 kW6U rack
NVIDIA DGX H2008× H200 SXM51.1 TB HBM3eNVLink 4.0 (900 GB/s)10.2 kW6U rack
NVIDIA HGX H1008× H100 SXM5640 GB HBM3NVLink 4.010.2 kWOEM chassis
OEM 4× GPU Server4× H100/H200 PCIe320 GB HBM3NVLink 3.0 (600 GB/s)5–6 kW2U rack

Network Fabric

The network fabric connects GPU nodes for distributed training and connects the cluster to storage and management networks. The fabric must provide sufficient bandwidth to prevent GPU idle time during all-reduce operations, the collective communication pattern used in distributed training.

AI Network Fabric Options

FabricBandwidthLatencyScaleCostBest For
InfiniBand NDR400 Gb/s per port~500 ns1,000+ GPUsHighLarge training clusters, frontier models
InfiniBand HDR200 Gb/s per port~600 ns500+ GPUsMedium-highMid-scale training, HPC
Ethernet 400 GbE + RoCEv2400 Gb/s per port1–3 µsUnlimitedMediumLarge inference, cost-sensitive training
Ethernet 100 GbE + RoCEv2100 Gb/s per port1–5 µsUnlimitedLow-mediumSmall clusters, inference, dev/test

Lossless fabric requirement

RoCEv2 requires a lossless Ethernet fabric, Priority Flow Control (PFC) and DCQCN congestion control must be correctly configured. Misconfigured RoCEv2 networks perform worse than standard TCP/IP. InfiniBand is inherently lossless and does not require this configuration.

Storage Architecture

AI training requires storage that can feed data to GPUs faster than they can consume it. A single DGX H100 system can consume data at 200+ GB/s during training. A cluster of 64 DGX systems requires storage capable of delivering 12+ TB/s of aggregate throughput, far beyond what traditional enterprise storage can provide.

Hot tier, Parallel file system

Technologies: GPFS, Lustre, WEKA, VAST, DDN

Active training datasets, checkpoints, model weights. Must deliver hundreds of GB/s of sequential throughput.

Warm tier, High-capacity NVMe

Technologies: All-NVMe arrays, NVMe-oF

Preprocessed datasets, recent checkpoints, inference model storage.

Cold tier, Object storage

Technologies: S3-compatible, Ceph, MinIO

Raw datasets, archived models, long-term checkpoint storage.

Power Engineering

Power engineering for AI infrastructure requires sizing for actual GPU utilization under training load, not nameplate ratings. A DGX H100 system draws 10.2 kW at full load. A rack of 4 DGX systems draws 40+ kW: requiring 60A 3-phase circuits, high-density PDUs, and UPS systems sized for the full cluster.

Power capacity is the most common constraint

Most existing data centers cannot support AI rack densities without significant power infrastructure upgrades. Assess available power capacity before committing to an AI infrastructure deployment in an existing facility.

Cooling Systems

Air cooling cannot remove heat from racks drawing 20+ kW at the densities AI clusters require. The three liquid cooling approaches for AI infrastructure are rear-door heat exchangers (RDHx), direct-to-chip liquid cooling (DLC), and full immersion cooling.

Cooling Technology Comparison

TechnologyMax Rack DensityCooling EfficiencyInfrastructure ChangeCost
Air coolingUp to 15 kWPUE 1.4–2.0NoneLow
Rear-door heat exchangerUp to 30 kWPUE 1.2–1.5Chilled water to rackMedium
Direct-to-chip liquidUp to 60 kWPUE 1.1–1.3Liquid loop to each serverHigh
Full immersionUp to 100+ kWPUE 1.02–1.1Complete facility redesignVery high

Design Patterns

Pod-based scaling

Design the cluster as a set of identical pods: each pod containing a fixed number of GPU nodes, network switches, and storage nodes. Scale by adding pods, not individual components.

Separate training and inference fabrics

Training requires high-bandwidth, low-latency InfiniBand. Inference requires high-throughput, low-latency Ethernet. Separate fabrics prevent training traffic from affecting inference latency.

Out-of-band management network

A dedicated management network (BMC, IPMI, iDRAC) separate from the data fabric allows infrastructure management without affecting training workloads.

Checkpoint storage co-location

Place checkpoint storage (NVMe) physically close to compute nodes to minimize checkpoint write latency, a critical factor for training job recovery time.

More AI Infrastructure Guides

Foundational

Beginner Overview

Plain-language introduction — what it is, why it matters, and how it fits into the broader infrastructure picture.

Strategic

Executive Brief

Business case, risk exposure, investment framing, and the three questions every executive should ask before approving a project.

Decision

Buying Guide

Vendor evaluation criteria, RFP requirements, contract terms to negotiate, and the questions that separate qualified vendors from unqualified ones.

Implementation

Planning Checklist

Pre-project checklist covering site readiness, stakeholder alignment, compliance requirements, and the decisions that must be made before work begins.

Strategic

Common Mistakes

The ten most expensive mistakes organizations make — and the specific decisions that prevent each one.

Foundational

Frequently Asked Questions

Direct answers to the questions procurement teams, IT leaders, and executives ask most often.

Implementation

Implementation Roadmap

Phase-by-phase delivery plan with milestones, dependencies, go/no-go criteria, and the decisions that determine schedule performance.

Decision

Comparison Guide

Side-by-side comparison of approaches, vendors, and architectures — with the criteria that matter for enterprise procurement decisions.

Strategic

Related Solutions

How this category connects to adjacent infrastructure domains — and the DCS Global solutions that address the full scope.

Decision

Recommended Next Steps

A decision tree for your specific situation — what to do next based on where you are in the planning or procurement process.

Related Categories

Apply This Knowledge

Ready to move from research to decision?

DCS Global engineers can review your specific requirements and give you a direct assessment, not a sales pitch. Our infrastructure specialists have delivered a broad portfolio of projects across North America, Europe, the Middle East, and Asia-Pacific.