Mistake 1: Deploying three-tier architecture for east-west traffic
Root Cause
Organizations that modernize compute and storage but retain the three-tier network architecture that was designed for north-south traffic.
Consequence
East-west traffic (server-to-server) must traverse multiple hops through the distribution and core layers, creating congestion and latency that limits application performance.
Prevention
Deploy spine-leaf architecture for all new data center network deployments. Three-tier architecture is inadequate for modern east-west traffic patterns.
Mistake 2: Undersizing for AI cluster bandwidth
Root Cause
Organizations that plan AI deployments without understanding the network bandwidth requirements, and deploy GPU servers on 10 GbE networks.
Consequence
GPU clusters waiting for data over congested networks have utilization rates of 20–40% instead of 80–90%. The cost of underutilized GPU infrastructure far exceeds the cost of adequate networking.
Prevention
Plan AI cluster networking before ordering GPU hardware. Verify that the network can deliver 200–400 Gb/s per node before deployment.
Mistake 3: Flat networks without segmentation
Root Cause
Organizations that deploy flat networks without VLAN segmentation or microsegmentation, because segmentation adds complexity.
Consequence
Attackers who compromise one system can move laterally to all other systems on the flat network. Compliance frameworks that require segmentation are not satisfied.
Prevention
Design network segmentation into the topology from the beginning. Retrofitting segmentation after deployment is more expensive and more disruptive than designing it in.
Mistake 4: No out-of-band management network
Root Cause
Organizations that manage network infrastructure through the production network, because a separate management network adds cost and complexity.
Consequence
When the production network fails, the management network fails with it, preventing remote troubleshooting and recovery. Network incidents that could be resolved remotely require on-site intervention.
Prevention
Deploy a dedicated out-of-band management network for all infrastructure components. The management network must be isolated from production traffic.
Mistake 5: Spanning tree in the data center
Root Cause
Organizations that retain spanning tree protocol in the data center because it is familiar, without understanding that it blocks 50% of available bandwidth.
Consequence
Spanning tree blocks redundant paths to prevent loops, wasting bandwidth and creating convergence delays when topology changes occur.
Prevention
Replace spanning tree with ECMP routing in the data center. Spine-leaf with BGP EVPN/VXLAN eliminates spanning tree and uses all available paths simultaneously.
Mistake 6: Insufficient buffer depth for AI workloads
Root Cause
Organizations that select switching hardware based on port count and price, without evaluating buffer depth for AI training workloads.
Consequence
Shallow-buffer switches drop packets during the bursty traffic that AI training generates, degrading training performance and increasing training time.
Prevention
Evaluate buffer depth when selecting switches for AI cluster networks. Deep-buffer switches (32+ MB) are required for AI training workloads.
Network mistakes compound over time