Why GPU Networking Is Different
GPU cluster networking is fundamentally different from conventional data center networking. In a conventional server cluster, network traffic is primarily client-server (north-south) — servers communicate with clients and storage. In a GPU training cluster, the dominant traffic pattern is all-to-all (east-west) — every GPU communicates with every other GPU during gradient synchronization. This all-reduce communication pattern requires high bandwidth, low latency, and RDMA (Remote Direct Memory Access) to minimize CPU overhead.
InfiniBand NDR
IB Latency
All-Reduce Impact
RoCEv2 Latency
InfiniBand for AI Clusters
InfiniBand is the dominant interconnect for large-scale AI training clusters. It provides native RDMA (bypassing the CPU for data transfers), sub-microsecond latency, and the highest available bandwidth. NVIDIA's acquisition of Mellanox (the primary InfiniBand vendor) has tightly integrated InfiniBand with NVIDIA GPU infrastructure.
Ethernet for AI Workloads
Ethernet is the dominant networking technology for inference clusters and is increasingly used for training clusters where InfiniBand's cost premium is not justified. Modern 400GbE with RoCEv2 provides near-InfiniBand performance at lower cost, with the advantage of a broader ecosystem and simpler integration with existing network infrastructure.
RoCE (RDMA over Converged Ethernet)
RoCEv2 enables RDMA over standard Ethernet infrastructure, providing near-InfiniBand performance at lower cost. However, RoCE requires careful network configuration to achieve low latency — standard Ethernet congestion control mechanisms are insufficient for RDMA traffic.
Network Topology for GPU Clusters
InfiniBand vs. Ethernet Comparison
GPU Cluster Networking Technology Comparison
| Technology | Bandwidth | Latency | RDMA | Cost | Ecosystem | Best For |
|---|---|---|---|---|---|---|
| InfiniBand HDR (200Gb) | 200 Gb/s per port | <1 µs | Native | High | NVIDIA/Mellanox | Large-scale training clusters |
| InfiniBand NDR (400Gb) | 400 Gb/s per port | <1 µs | Native | Very high | NVIDIA/Mellanox | Hyperscale training, frontier models |
| RoCEv2 (100GbE) | 100 Gb/s per port | 1–5 µs | Yes (with ECN) | Medium | Broad | Mid-scale training, inference clusters |
| RoCEv2 (400GbE) | 400 Gb/s per port | 1–5 µs | Yes (with ECN) | High | Broad | Large-scale training alternative to IB |
| Standard Ethernet (100GbE) | 100 Gb/s per port | 10–100 µs | No | Low-Medium | Universal | Inference, management, storage |
| NVLink (intra-server) | 900 GB/s per GPU | <1 µs | N/A (GPU-GPU) | Included in SXM | NVIDIA only | Intra-server GPU communication |
Frequently Asked Questions
Do I need InfiniBand for my GPU cluster?
It depends on your workload. For large-scale training (16+ GPUs, large models): InfiniBand HDR or NDR is strongly recommended. The all-reduce communication overhead with standard Ethernet can reduce training throughput by 20–40% compared to InfiniBand. For inference clusters: standard Ethernet (100GbE or 400GbE) is sufficient — inference workloads do not require all-reduce communication. For small training clusters (8 GPUs within a single DGX): NVLink handles intra-server communication; InfiniBand is only needed for inter-server communication. For mid-scale training (16–64 GPUs): RoCEv2 with 400GbE is a cost-effective alternative to InfiniBand with 10–20% performance penalty.
What is the cost difference between InfiniBand and Ethernet for a GPU cluster?
InfiniBand infrastructure costs 2–3× more than equivalent Ethernet. For a 16-node DGX H100 cluster: InfiniBand HDR fabric (switches + cables + NICs): ~$200K–$400K. Equivalent 400GbE Ethernet fabric: ~$80K–$150K. The performance premium of InfiniBand (10–20% higher training throughput) must be weighed against the cost premium. For a $5M GPU cluster, the $200K networking premium is ~4% of total cost — often justified. For smaller deployments, RoCEv2 Ethernet may be more cost-effective.
How many network ports does each GPU server need?
DGX H100: 8× HDR200 InfiniBand ports (one per GPU) + 2× 10GbE management ports + 2× 100GbE storage/data ports. HGX H100 8-GPU: similar configuration. For inference servers with PCIe GPUs: 2× 100GbE or 400GbE for inference traffic + 1× 10GbE management. The key principle: each GPU should have its own dedicated network port for training workloads — sharing ports between GPUs creates network bottlenecks during all-reduce operations.