Spine-Leaf Design Principles
Spine-leaf topology consists of two layers: leaf switches (connected to servers, storage, and other endpoints) and spine switches (connected only to leaf switches). Every leaf connects to every spine: creating a non-blocking, any-to-any topology where every server is exactly two hops from every other server.
Oversubscription ratio
The ratio of downlink bandwidth (to servers) to uplink bandwidth (to spines). A 3:1 oversubscription ratio means 3 Gb/s of downlink for every 1 Gb/s of uplink. Lower oversubscription provides better performance for east-west traffic.
ECMP load balancing
Equal-Cost Multi-Path routing distributes traffic across all available paths simultaneously. A leaf switch with 4 uplinks to 4 spine switches uses all 4 paths, providing 4x the bandwidth of a single path.
Leaf switch sizing
Leaf switches are sized for the number of servers they connect and the uplink bandwidth required. A typical leaf switch has 48 downlinks (25/100 GbE) and 8 uplinks (100/400 GbE).
Spine switch sizing
Spine switches are sized for the number of leaf switches they connect. A spine switch with 32 ports at 400 GbE can connect 32 leaf switches with 400 Gb/s uplinks each.
Overlay Networking: BGP EVPN with VXLAN
BGP EVPN (Ethernet VPN) with VXLAN (Virtual Extensible LAN) is the standard overlay networking technology for modern data centers. VXLAN encapsulates Layer 2 frames in UDP packets, enabling Layer 2 connectivity across a Layer 3 underlay network. BGP EVPN provides the control plane for VXLAN, distributing MAC and IP address information across the fabric without flooding.
Why VXLAN replaced spanning tree
AI Cluster Interconnects
AI Cluster Interconnect Comparison
| Technology | Bandwidth | Latency | Protocol | Best For |
|---|---|---|---|---|
| InfiniBand HDR | 200 Gb/s per port | ~600 nanoseconds | RDMA | Large AI training clusters, HPC |
| InfiniBand NDR | 400 Gb/s per port | ~600 nanoseconds | RDMA | Largest AI clusters, highest performance |
| RoCEv2 (100 GbE) | 100 Gb/s per port | ~1 microsecond | RDMA over Ethernet | AI clusters on existing Ethernet infrastructure |
| RoCEv2 (400 GbE) | 400 Gb/s per port | ~1 microsecond | RDMA over Ethernet | High-performance AI on Ethernet |
| NVLink (within server) | 900 GB/s total | ~1 microsecond | Proprietary | GPU-to-GPU within a single server |
InfiniBand provides lower latency than Ethernet for AI training workloads. RoCEv2 provides comparable bandwidth on existing Ethernet infrastructure, but requires lossless Ethernet configuration (PFC + ECN) to achieve near-InfiniBand performance.
Network Security Architecture
Microsegmentation
Workload-level network isolation using distributed firewalling. Policies follow workloads, not network segments. Limits lateral movement to individual workloads rather than network segments.
Zero-trust network access
Every access request is verified regardless of network location. Identity-based access policies replace network-location-based trust.
Encrypted east-west traffic
MACsec (Layer 2) or IPsec (Layer 3) encryption for traffic between servers. Required for compliance frameworks that mandate encryption in transit.
Network telemetry
Streaming telemetry from all network devices to a security information and event management (SIEM) platform. Provides visibility into traffic patterns and anomalies.
Network Automation
Network automation reduces configuration errors, accelerates change deployment, and enables infrastructure-as-code workflows. Modern data center networks are configured through APIs: not CLI commands: enabling version-controlled, tested, and repeatable configuration management.
Key automation tools include Ansible (agentless configuration management), Terraform (infrastructure-as-code provisioning), and NAPALM (network automation and programmability abstraction layer). Networks that are not automated are more expensive to operate and more prone to configuration errors.
Network OEM Comparison
Data Center Switching OEM Comparison
| OEM | Strengths | AI/HPC Capability | Automation |
|---|---|---|---|
| Cisco (Nexus) | Broad enterprise support, ACI SDN, strong ISV ecosystem | Strong (Nexus 9000) | Excellent (NSO, ACI) |
| Arista Networks | EOS consistency, CloudVision, strong automation | Very strong (7800R3) | Excellent (CloudVision, eAPI) |
| NVIDIA (Spectrum) | InfiniBand + Ethernet convergence, AI-optimized | Purpose-built (Quantum-2) | Good (NVIDIA Air) |
| Juniper (QFX) | Apstra intent-based, strong automation | Strong (QFX10000) | Excellent (Apstra) |
| HPE (Aruba) | Fabric Composer, strong campus integration | Adequate | Good (Fabric Composer) |