Planning Methodology
Network capacity planning follows a four-step process:
- Baseline: Measure current traffic patterns, utilization, and port density
- Forecast: Project future requirements based on business growth, new workloads, and technology changes
- Gap analysis: Identify where current capacity will be insufficient for projected requirements
- Plan: Define upgrade actions, timelines, and budgets to close identified gaps
Capacity planning should look 18–24 months ahead. This horizon accounts for: equipment procurement lead times (6–12 months for some network hardware), installation and testing time, and budget planning cycles. Capacity constraints discovered with less than 6 months lead time often cannot be resolved before they impact operations.
Traffic Analysis
Traffic Measurement
Accurate capacity planning requires accurate traffic measurement. Key metrics:
- Average utilization: Average bandwidth utilization over a measurement period (typically 5-minute intervals)
- Peak utilization: Maximum utilization observed during the measurement period
- 95th percentile utilization: The utilization level exceeded only 5% of the time — the standard metric for capacity planning
- Traffic direction: Ratio of east-west to north-south traffic
- Traffic patterns: Time-of-day and day-of-week patterns
Traffic Sources
Collect traffic data from: SNMP polling of switch interfaces (5-minute intervals minimum), NetFlow/IPFIX for traffic flow analysis, and application performance monitoring for application-level traffic patterns.
Traffic Growth Trends
Analyze historical traffic trends to project future growth. Data center east-west traffic has grown 30–50% annually in recent years, driven by microservices, distributed databases, and AI workloads. Apply growth rates to current measurements to project future requirements.
Bandwidth Modeling
Utilization Targets
Target utilization levels for capacity planning:
- Core/spine links: 60–70% average utilization; upgrade when 95th percentile exceeds 80%
- Leaf uplinks: 50–60% average utilization; upgrade when 95th percentile exceeds 75%
- Server access links: 40–50% average utilization; upgrade when 95th percentile exceeds 70%
- WAN/internet links: 50–60% average utilization; upgrade when 95th percentile exceeds 75%
Oversubscription Modeling
Model oversubscription ratios based on actual traffic patterns. If servers generate 10 Gbps of traffic but only 2 Gbps is destined for other servers (east-west), a 5:1 oversubscription ratio may be acceptable. If 8 Gbps is east-west traffic, a 1.25:1 oversubscription ratio is required.
Burst Capacity
Plan for traffic bursts that exceed average utilization. Batch jobs, backup windows, and application deployments can cause traffic spikes 3–5x average utilization. Ensure sufficient headroom for expected burst patterns.
Port Density Planning
Port density planning ensures that sufficient switch ports are available for current and future server deployments.
Current Inventory
Document current port utilization: total ports, used ports, and available ports by switch and by rack. Identify switches approaching port exhaustion.
Growth Projection
Project future port requirements based on: planned server deployments, storage system additions, and network appliance additions. Include ports for management, out-of-band, and redundant connections.
Port Speed Upgrades
Plan for port speed upgrades as server NIC speeds increase. The industry is transitioning from 25GbE to 100GbE server connections; AI servers require 400GbE. Leaf switches must support the port speeds required by current and future servers.
AI Workload Capacity
AI training workloads require fundamentally different capacity planning than traditional workloads:
All-Reduce Traffic
During all-reduce operations, every GPU communicates with every other GPU simultaneously. A 64-GPU cluster generates 64 × 400 Gb/s = 25.6 Tb/s of all-reduce traffic. The network fabric must support this traffic without congestion.
Non-Blocking Requirement
AI training requires non-blocking (1:1 oversubscription) fabric. Traditional capacity planning with 2:1 or 4:1 oversubscription is insufficient — any congestion causes GPU stalls and reduces training efficiency.
Storage Bandwidth
Training data must be delivered to GPUs at sufficient throughput to keep them fed. Plan storage network bandwidth separately from compute fabric: 5–10 GB/s per 8-GPU server for typical training workloads.
Monitoring & Alerting
Capacity planning is only as good as the monitoring data it is based on. Key monitoring requirements:
- Interface utilization: 5-minute polling of all switch interfaces via SNMP or streaming telemetry
- Error counters: Monitor for interface errors, discards, and CRC errors that indicate physical layer problems
- Queue depth: Monitor switch queue depths for congestion indicators
- Flow data: NetFlow/IPFIX for traffic flow analysis and anomaly detection
Alerting Thresholds
Set alerts at: 70% utilization (warning), 85% utilization (critical). Alerts should trigger capacity review, not emergency action — capacity planning should prevent reaching critical thresholds.
Upgrade Planning
Upgrade Triggers
Initiate upgrade planning when: 95th percentile utilization exceeds 75% on critical links, port utilization exceeds 80%, or new workloads are planned that will significantly increase traffic.
Upgrade Options
- Link aggregation: Bond multiple links for higher bandwidth (limited by switch port availability)
- Port speed upgrade: Replace 10GbE with 25GbE, or 100GbE with 400GbE
- Additional spine switches: Add spine switches to increase bandwidth between leaf switches
- Additional leaf switches: Add leaf switches to add server capacity
- Full fabric replacement: Replace the entire fabric for major architecture changes
Upgrade Execution
Network upgrades in live environments require careful planning: maintenance windows, rollback procedures, and staged rollouts. Test upgrades in non-production environments before applying to production. Document all changes and verify performance after each change.