Architecture Questions
Q: What is spine-leaf architecture and why is it better than three-tier?
Spine-leaf architecture uses two tiers: leaf switches (connecting servers) and spine switches (connecting leaf switches). Every leaf connects to every spine. This provides predictable two-hop latency between any two servers, full bandwidth utilization through ECMP, and easy horizontal scaling. Three-tier (core-distribution-access) was designed for north-south traffic; spine-leaf is optimized for the east-west traffic that dominates modern data centers.
Q: What is oversubscription and what ratio should I use?
Oversubscription is the ratio of server-facing bandwidth to spine-facing bandwidth on a leaf switch. A 2:1 oversubscription means servers can generate twice the bandwidth that the uplinks to spine can carry. For most enterprise workloads, 2:1 to 4:1 is acceptable. For AI training workloads, 1:1 (non-blocking) is required — any congestion causes GPU stalls.
Q: What is ECMP and why does it matter?
Equal-Cost Multi-Path routing distributes traffic across multiple equal-cost paths. In a spine-leaf network with 4 spine switches, ECMP distributes traffic across all 4 paths, providing 4x the bandwidth of a single path. Without ECMP, only one path would be used and the others would be wasted. ECMP is the mechanism that makes spine-leaf networks highly efficient.
Q: What is VXLAN and why is it used in data centers?
VXLAN (Virtual Extensible LAN) encapsulates Layer 2 Ethernet frames in UDP packets, enabling Layer 2 networks to span Layer 3 boundaries. This allows virtual machines and containers to maintain their IP addresses as they move between physical hosts. VXLAN with EVPN control plane is the standard overlay for modern spine-leaf networks.
Q: Why is BGP used in data center networks instead of OSPF?
BGP provides better scalability, more flexible routing policy control, and native ECMP support compared to OSPF for large spine-leaf networks. BGP also avoids the convergence issues that OSPF can experience in large networks. Most modern spine-leaf implementations use eBGP between leaf and spine switches.
AI Networking Questions
Q: What network bandwidth does an AI training cluster need?
NVIDIA recommends 400 Gb/s per GPU for optimal H100/H200 training efficiency. A 64-GPU cluster requires 64 × 400 Gb/s = 25.6 Tb/s of aggregate fabric bandwidth. The network must deliver this bandwidth without congestion — any bottleneck reduces GPU utilization.
Q: InfiniBand or Ethernet for AI clusters?
InfiniBand NDR (400 Gb/s) provides the best all-reduce performance (1–2 μs latency, native RDMA, NVIDIA SHARP) but at higher cost and with NVIDIA-only vendor lock-in. RoCEv2 over 400GbE provides lower cost and vendor diversity but higher latency (2–5 μs) and requires careful PFC/ECN configuration. For maximum training efficiency, InfiniBand is preferred. For cost-sensitive deployments, RoCEv2 is viable.
Q: What is RDMA and why does AI need it?
Remote Direct Memory Access enables GPU-to-GPU data transfer without CPU involvement. During all-reduce operations, GPUs exchange gradient data directly over the network without the CPU copying data between GPU memory and network buffers. This reduces latency and CPU overhead significantly. InfiniBand provides native RDMA; RoCEv2 provides RDMA over Ethernet.
Q: What is NVIDIA SHARP?
Scalable Hierarchical Aggregation and Reduction Protocol — NVIDIA's in-network computing technology for InfiniBand. SHARP performs all-reduce operations within the switch fabric rather than at the endpoints, reducing the amount of data that must traverse the network. Can improve all-reduce performance by 2x for large clusters.
Q: Why is the network often the bottleneck in AI training?
During distributed training, GPUs must synchronize gradients after each training step. If the network cannot deliver gradient data fast enough, GPUs stall waiting for synchronization. A 64-GPU cluster with insufficient network bandwidth can have effective GPU utilization of 40–60% instead of 90%+. Network fabric design is often the most underestimated component of AI infrastructure.
Protocol Questions
Q: What is the difference between Layer 2 and Layer 3 switching?
Layer 2 switching forwards traffic based on MAC addresses within a single network segment. Layer 3 switching (routing) forwards traffic based on IP addresses between network segments. Modern spine-leaf networks use Layer 3 at the leaf switch — servers connect at Layer 2, but inter-leaf communication is routed at Layer 3.
Q: What is PFC and why is it required for RoCEv2?
Priority Flow Control is an Ethernet flow control mechanism that pauses specific traffic classes to prevent packet loss. RoCEv2 requires lossless Ethernet because RDMA does not handle packet loss gracefully — a single dropped packet causes significant performance degradation. PFC prevents packet loss by pausing traffic before buffers overflow.
Q: What is SD-WAN and when should I use it?
SD-WAN uses software to manage WAN connectivity across multiple transport types (MPLS, broadband, LTE). It provides application-aware routing, automatic failover, and centralized management. Use SD-WAN for: branch office connectivity (replace expensive MPLS with broadband), cloud connectivity optimization, and multi-site WAN management. SD-WAN is not a replacement for dedicated connectivity (Direct Connect/ExpressRoute) to data centers.
Cabling Questions
Q: What fiber type should I use for 400GbE?
For distances up to 100m: OM4 or OM5 multimode fiber with QSFP-DD SR8 transceivers. For distances up to 150m: OM5 multimode fiber. For distances beyond 150m: OS2 single-mode fiber with QSFP-DD LR8 transceivers. For within-rack connections (up to 3m): QSFP-DD DAC (Direct Attach Copper) cables are the lowest-cost option.
Q: What is the difference between DAC and AOC cables?
DAC (Direct Attach Copper) cables are copper cables with integrated transceivers at each end. Passive DAC works up to 3–5m; active DAC works up to 7–15m. AOC (Active Optical Cable) cables use fiber optics with integrated transceivers. AOC works up to 100m and provides better signal integrity than long DAC cables. DAC is lower cost for short distances; AOC is required for longer distances.
Q: What cabling standard should I follow for my data center?
Follow ANSI/TIA-942 for data center cabling infrastructure and ANSI/TIA-568 for cabling specifications. TIA-942 defines topology requirements, cabling distances, and redundancy requirements. TIA-568 specifies cable categories, connector types, and performance requirements. For international deployments, ISO/IEC 11801 is the equivalent standard.
Operations Questions
Q: How do I monitor network utilization?
Use SNMP polling (5-minute intervals) for interface utilization metrics. For more granular data, use streaming telemetry (gNMI/gRPC) for sub-minute metrics. NetFlow/IPFIX provides traffic flow analysis. Target 60–70% utilization on critical links; alert at 80% utilization.
Q: How often should I upgrade my network infrastructure?
Network hardware typically has a 5–7 year lifecycle. Plan upgrades based on: capacity (when utilization consistently exceeds 70–75%), port speed (when servers require faster connections than available), and end-of-support (when vendor support ends). For AI infrastructure, upgrade cycles may be shorter due to rapid GPU hardware evolution.
Q: What is zero-touch provisioning?
Zero-touch provisioning (ZTP) enables network switches to automatically download and apply their configuration when first powered on, without manual intervention. The switch boots, contacts a DHCP server to get its IP address and the location of a configuration server, downloads its configuration, and applies it. ZTP dramatically reduces the time and expertise required to deploy new network equipment.