Design Principles

Spine-leaf architecture is based on the Clos network topology, originally designed for telephone switching in the 1950s and adapted for data center networking. The key design principles:

Two-Tier Hierarchy

Leaf switches connect to servers and other endpoints. Spine switches connect only to leaf switches. Every leaf connects to every spine; spines do not connect to each other. This creates a non-blocking, full-mesh connectivity between all leaf switches through the spine layer.

Layer 3 at the Leaf

Modern spine-leaf networks run Layer 3 (IP routing) at the leaf switch, eliminating Spanning Tree Protocol and enabling ECMP. Each leaf switch is a Layer 3 boundary — servers connect to the leaf at Layer 2, but inter-leaf communication is routed at Layer 3.

Symmetric Design

All leaf switches have the same number of uplinks to spine switches. All spine switches have the same number of downlinks to leaf switches. This symmetry ensures that ECMP distributes traffic evenly across all paths.

Sizing & Oversubscription

Oversubscription Ratio

Oversubscription is the ratio of server-facing bandwidth to spine-facing bandwidth on a leaf switch. A leaf switch with 48x 25GbE server ports (1,200 Gbps total) and 8x 100GbE uplinks to spine (800 Gbps total) has an oversubscription ratio of 1.5:1.

Common oversubscription ratios:

  • 1:1 (non-blocking): Full bandwidth between any two servers. Required for AI training workloads. Most expensive.
  • 2:1: Appropriate for most enterprise workloads. Good balance of cost and performance.
  • 4:1: Acceptable for workloads with bursty traffic patterns. Lower cost.
  • 8:1 or higher: Only appropriate for workloads with very low bandwidth requirements.

Leaf Switch Sizing

Size leaf switches based on: number of servers per rack row, server port speed (10GbE, 25GbE, 100GbE), required oversubscription ratio, and number of spine uplinks. Common leaf switch configurations: 48x 25GbE + 8x 100GbE uplinks (2:1 oversubscription), or 32x 100GbE + 8x 400GbE uplinks (1:1 for AI workloads).

Spine Switch Sizing

Size spine switches based on: number of leaf switches, uplink speed from leaf switches, and required oversubscription. A spine switch must have enough ports to connect to all leaf switches. Common spine switch configurations: 32x 400GbE or 64x 400GbE for large fabrics.

BGP in Spine-Leaf Networks

BGP (Border Gateway Protocol) has become the standard routing protocol for spine-leaf networks, replacing OSPF and EIGRP used in traditional three-tier networks. BGP provides:

  • Scalability: BGP scales to very large networks without the convergence issues of link-state protocols
  • ECMP support: BGP supports ECMP across multiple equal-cost paths
  • Policy control: BGP provides fine-grained routing policy control
  • Vendor interoperability: BGP is a standard protocol supported by all network vendors

eBGP vs. iBGP

Most spine-leaf implementations use eBGP (external BGP) between leaf and spine switches, with each switch in its own BGP autonomous system (AS). This simplifies configuration and avoids the full-mesh requirement of iBGP. Each leaf switch has a unique AS number; spine switches share an AS number or each have unique AS numbers.

VXLAN/EVPN Overlay

VXLAN (Virtual Extensible LAN) with EVPN (Ethernet VPN) control plane is the standard overlay for spine-leaf networks. It enables:

  • Layer 2 connectivity between servers on different leaf switches
  • VM and container mobility without IP address changes
  • Network virtualization for multi-tenant environments
  • Efficient handling of broadcast and multicast traffic

EVPN uses BGP to distribute MAC and IP address information between leaf switches (VTEPs — VXLAN Tunnel Endpoints). This eliminates the need for flood-and-learn MAC address discovery, reducing unnecessary traffic in the fabric.

AI Workload Design

AI training workloads have specific spine-leaf design requirements:

Non-Blocking Fabric

AI all-reduce operations require every GPU to communicate with every other GPU simultaneously. Any congestion in the fabric reduces GPU utilization. AI clusters require non-blocking (1:1 oversubscription) fabric design.

High-Speed Leaf Switches

AI servers with 8x H100 GPUs require 400GbE or InfiniBand NDR uplinks from the leaf switch. Leaf switches for AI clusters must support 400GbE or 800GbE ports.

Fat-Tree Topology

For large AI clusters, a fat-tree topology (multiple tiers of spine switches) provides the bandwidth required for all-reduce operations at scale. NVIDIA Quantum-2 InfiniBand switches support fat-tree topologies for clusters of up to 32,768 GPUs.

RDMA Support

AI clusters require RDMA (Remote Direct Memory Access) support in the network fabric. For Ethernet-based AI fabrics, RoCEv2 (RDMA over Converged Ethernet) requires Priority Flow Control (PFC) and Explicit Congestion Notification (ECN) configuration on all switches.

Scaling the Fabric

Adding Leaf Switches

Add leaf switches to add server capacity. Each new leaf switch connects to all existing spine switches. No changes required to existing leaf or spine switches. This is the primary scaling mechanism for spine-leaf networks.

Adding Spine Switches

Add spine switches to add bandwidth between leaf switches. Each new spine switch connects to all existing leaf switches. Requires adding uplink ports to all leaf switches — plan leaf switch port density for future spine expansion.

Multi-Tier Spine-Leaf

For very large data centers, a multi-tier spine-leaf (super-spine layer) extends the architecture. Super-spine switches connect multiple spine-leaf pods, enabling data centers with thousands of leaf switches. Used by hyperscale cloud providers and large enterprise data centers.

Vendor Options

VendorPlatformKey Strengths
CiscoNexus 9000 / ACILargest installed base; ACI SDN controller; strong enterprise support
Arista7000 Series / EOSProgrammable EOS; strong in financial services and cloud; excellent automation
JuniperQFX Series / ApstraApstra intent-based networking; strong in service provider environments
NVIDIASpectrum / Quantum-2Best-in-class for AI workloads; InfiniBand NDR; Spectrum-X for Ethernet AI
BroadcomTomahawk / TridentDominant ASIC vendor; used by many white-box and OEM switch vendors