Networking Fundamentals
Data center networking operates across multiple layers of the OSI model:
- Physical layer (Layer 1): Cables, transceivers, and physical ports. Copper (DAC cables for short distances) and fiber optic (for longer distances and higher speeds).
- Data link layer (Layer 2): Ethernet frames, MAC addresses, VLANs. Switches operate at this layer.
- Network layer (Layer 3): IP addresses, routing. Routers and Layer 3 switches operate at this layer.
- Transport layer (Layer 4): TCP/UDP. Relevant for load balancing and firewall rules.
Modern data center networks increasingly push routing to the edge (leaf switches) and use BGP (Border Gateway Protocol) as the routing protocol within the data center — a significant departure from traditional enterprise networking that used OSPF or EIGRP.
Traditional Three-Tier Architecture
The traditional data center network used a three-tier hierarchy:
- Access layer: Switches connecting servers to the network
- Distribution layer: Aggregation switches connecting access switches
- Core layer: High-speed switches providing connectivity between distribution layers and to the WAN
Three-tier architecture was designed for north-south traffic (client-to-server). It has significant limitations for modern data centers: traffic between servers in different access switches must traverse multiple hops, creating latency and bottlenecks. Spanning Tree Protocol (STP) blocked redundant paths, wasting bandwidth. Scaling required adding distribution and core switches, increasing complexity.
Spine-Leaf Architecture
Spine-leaf (also called Clos network) architecture addresses the limitations of three-tier networks:
- Leaf switches: Connect to servers, storage, and other endpoints. Every leaf switch connects to every spine switch.
- Spine switches: Connect only to leaf switches — no server connections. Every spine switch connects to every leaf switch.
Key properties of spine-leaf architecture:
- Predictable latency: Any server can reach any other server in exactly two hops (leaf → spine → leaf)
- Full bandwidth utilization: ECMP routing distributes traffic across all spine switches, utilizing all available bandwidth
- Easy scaling: Add leaf switches to add server capacity; add spine switches to add bandwidth
- No spanning tree: Layer 3 routing at the leaf eliminates STP and its blocked paths
ECMP (Equal-Cost Multi-Path)
ECMP routing distributes traffic across multiple equal-cost paths. In a spine-leaf network with 4 spine switches, traffic from any leaf to any other leaf can use any of the 4 paths — providing 4x the bandwidth of a single path. ECMP is the mechanism that makes spine-leaf networks highly efficient.
Traffic Patterns
Modern data center traffic patterns have fundamentally changed:
East-West Traffic Dominance
Traditional data centers had predominantly north-south traffic (client requests to servers). Modern data centers have predominantly east-west traffic (server-to-server): microservices calling each other, distributed databases replicating, storage systems communicating, and AI training nodes exchanging gradients.
East-west traffic can represent 70–80% of total data center traffic in modern environments. Three-tier architectures were not designed for this traffic pattern; spine-leaf architectures are.
AI Training Traffic
AI training workloads generate extremely high east-west traffic: all-reduce operations during distributed training require every GPU to communicate with every other GPU simultaneously. This creates a traffic pattern that requires non-blocking, high-bandwidth fabrics — the primary driver of InfiniBand and 400GbE adoption.
Overlay Networks
Overlay networks create virtual networks on top of the physical network, enabling network virtualization independent of physical topology.
VXLAN (Virtual Extensible LAN)
VXLAN encapsulates Layer 2 Ethernet frames in UDP packets, enabling Layer 2 networks to span Layer 3 boundaries. This allows virtual machines and containers to maintain their IP addresses as they move between physical hosts. VXLAN is the standard overlay technology for modern data center networks.
EVPN (Ethernet VPN)
EVPN is a BGP-based control plane for VXLAN overlays. It provides: MAC/IP address learning and distribution, multi-homing support, and efficient handling of broadcast, unknown unicast, and multicast (BUM) traffic. EVPN + VXLAN is the standard for modern data center overlay networks.
Software-Defined Networking
SDN separates the network control plane (decisions about where traffic goes) from the data plane (forwarding traffic). A centralized SDN controller manages the network programmatically, enabling automation and dynamic reconfiguration.
SDN benefits: automated network provisioning, consistent policy enforcement, rapid response to changing traffic patterns, and simplified operations. SDN is increasingly used in large data centers and cloud environments where manual network management is impractical.
Common SDN implementations: Cisco ACI (Application Centric Infrastructure), VMware NSX, and open-source platforms (OpenDaylight, ONOS).
AI Networking Requirements
AI training workloads have networking requirements that differ significantly from traditional enterprise workloads:
- Bandwidth: All-reduce operations require every GPU to communicate with every other GPU simultaneously. A 64-GPU cluster requires 400 Gb/s per GPU for efficient training.
- Latency: All-reduce latency directly affects training throughput. InfiniBand NDR provides lower latency than Ethernet for all-reduce operations.
- Non-blocking fabric: Any congestion in the network fabric reduces GPU utilization. Non-blocking fat-tree topologies ensure full bandwidth between any two nodes.
- RDMA: Remote Direct Memory Access enables GPU-to-GPU data transfer without CPU involvement, reducing latency and CPU overhead.
Two technologies dominate AI cluster networking: InfiniBand NDR (400 Gb/s, NVIDIA Quantum-2 switches) and RoCEv2 over 400GbE. InfiniBand provides lower latency and better all-reduce performance; RoCEv2 provides lower cost and compatibility with standard Ethernet infrastructure.