AI Fabric Requirements

AI training workloads generate a specific traffic pattern: all-reduce operations require every GPU to exchange gradient data with every other GPU simultaneously. This creates a many-to-many communication pattern that is fundamentally different from traditional client-server traffic.

Bandwidth Requirements

Each GPU requires sufficient network bandwidth to avoid becoming a bottleneck during all-reduce. For H100 SXM5 GPUs (3.35 TB/s HBM3 bandwidth), the network must deliver at least 400 Gb/s per GPU to avoid being the bottleneck. NVIDIA recommends 400 Gb/s per GPU for optimal training efficiency.

Latency Requirements

All-reduce latency directly affects training throughput. Lower latency means faster gradient synchronization and higher GPU utilization. InfiniBand provides 1–2 microsecond latency; RoCEv2 provides 2–5 microsecond latency. For large clusters, the difference is significant.

Non-Blocking Fabric

Any congestion in the network fabric causes GPU stalls — GPUs wait for gradient data instead of computing. Non-blocking (1:1 oversubscription) fabric design is required for maximum training efficiency. Blocking fabrics can reduce effective GPU utilization to 40–60% of theoretical maximum.

InfiniBand NDR

NVIDIA Quantum-2 InfiniBand NDR (Next Data Rate) provides 400 Gb/s per port with 1–2 microsecond latency. It is the gold standard for AI training fabrics.

Key Specifications

  • 400 Gb/s per port (NDR); 200 Gb/s per port (HDR)
  • 1–2 microsecond end-to-end latency
  • RDMA (Remote Direct Memory Access) native support
  • NVIDIA Quantum-2 switches: 64x NDR400 ports per switch
  • NVLink 4.0 within DGX H100 servers (900 GB/s bidirectional)

Advantages

  • Lowest latency for all-reduce operations
  • Native RDMA — no configuration required for RDMA operation
  • NVIDIA SHARP (Scalable Hierarchical Aggregation and Reduction Protocol) — in-network computing that reduces all-reduce traffic by performing aggregation in the switch
  • Best scaling efficiency for large clusters

Disadvantages

  • Higher cost than Ethernet (switches, cables, HCAs)
  • Proprietary ecosystem — limited to NVIDIA switches and compatible HCAs
  • Separate network from Ethernet — requires separate management and cabling

RoCEv2 over Ethernet

RoCEv2 (RDMA over Converged Ethernet version 2) enables RDMA over standard Ethernet infrastructure. It provides a lower-cost alternative to InfiniBand for AI cluster networking.

Key Specifications

  • 400 Gb/s per port (400GbE); 800 Gb/s per port (800GbE, emerging)
  • 2–5 microsecond end-to-end latency (higher than InfiniBand)
  • RDMA over standard Ethernet infrastructure
  • Compatible with standard Ethernet switches (Arista, Cisco, Juniper, NVIDIA Spectrum)

Advantages

  • Lower cost than InfiniBand
  • Uses standard Ethernet infrastructure — compatible with existing management tools
  • Vendor diversity — multiple switch vendors support RoCEv2
  • Converged network — can share infrastructure with storage and management traffic

Disadvantages

  • Higher latency than InfiniBand
  • Requires careful PFC and ECN configuration — misconfiguration causes performance degradation
  • No equivalent to NVIDIA SHARP for in-network computing
  • Performance gap vs. InfiniBand is most significant for large clusters

Topology Design

Single-Switch Topology (up to 64 GPUs)

For clusters up to 64 GPUs, a single top-of-rack switch (NVIDIA Quantum-2 for InfiniBand, or 400GbE switch for Ethernet) provides non-blocking connectivity. All servers connect to the single switch; no inter-switch traffic required.

Two-Tier Fat-Tree (64–512 GPUs)

For clusters of 64–512 GPUs, a two-tier fat-tree topology: leaf switches connect to servers; spine switches connect to all leaf switches. Non-blocking design requires equal numbers of uplinks and downlinks on each leaf switch.

Example: 256-GPU cluster (32 servers × 8 GPUs). 4 leaf switches × 8 servers each. Each leaf has 8 server ports + 8 spine uplinks. 8 spine switches × 4 leaf downlinks each. Result: non-blocking, 400 Gb/s per GPU.

Three-Tier Fat-Tree (512+ GPUs)

For clusters exceeding 512 GPUs, a three-tier fat-tree adds a super-spine layer. NVIDIA Quantum-2 supports fat-tree topologies for clusters up to 32,768 GPUs. Three-tier designs require careful planning to maintain non-blocking properties.

Storage Network

The storage network must be designed separately from the compute fabric. Training data must be delivered to GPUs fast enough to keep them fed — storage network bandwidth is often the hidden bottleneck in AI clusters.

Storage Network Requirements

A 64-GPU cluster training a large language model may require 400–800 GB/s aggregate storage throughput. This requires a dedicated high-speed storage network: 100GbE or 200GbE per storage node, with sufficient switch capacity to deliver aggregate throughput.

Separate vs. Converged

Dedicated storage network (separate from compute fabric) provides the best performance and isolation. Converged network (storage and compute on the same fabric) reduces cost and complexity but requires careful QoS configuration to prevent storage traffic from impacting compute traffic.

Configuration Requirements

RoCEv2 Configuration

RoCEv2 requires specific switch configuration to prevent congestion-induced performance degradation:

  • Priority Flow Control (PFC): Enables lossless Ethernet for RDMA traffic. Must be enabled on all switches and NICs in the fabric.
  • Explicit Congestion Notification (ECN): Signals congestion before packet loss occurs. Must be configured on all switches.
  • DSCP marking: Mark RDMA traffic with appropriate DSCP values for QoS treatment.
  • Buffer tuning: Configure switch buffers for RDMA traffic patterns.

InfiniBand Configuration

InfiniBand requires a Subnet Manager (SM) to manage the fabric. NVIDIA OpenSM is the standard open-source SM. The SM discovers the fabric topology, assigns LIDs (Local Identifiers), and computes routing tables. For large fabrics, a dedicated SM server is recommended.

Technology Selection

CriterionInfiniBand NDRRoCEv2 400GbE
Latency1–2 μs (best)2–5 μs
Bandwidth400 Gb/s per port400 Gb/s per port
CostHigherLower
Vendor diversityNVIDIA onlyMultiple vendors
Configuration complexityLower (native RDMA)Higher (PFC/ECN required)
In-network computingNVIDIA SHARPNot available
Best forLarge clusters, maximum performanceCost-sensitive, smaller clusters