Planning Methodology

Network capacity planning follows a four-step process:

  1. Baseline: Measure current traffic patterns, utilization, and port density
  2. Forecast: Project future requirements based on business growth, new workloads, and technology changes
  3. Gap analysis: Identify where current capacity will be insufficient for projected requirements
  4. Plan: Define upgrade actions, timelines, and budgets to close identified gaps

Capacity planning should look 18–24 months ahead. This horizon accounts for: equipment procurement lead times (6–12 months for some network hardware), installation and testing time, and budget planning cycles. Capacity constraints discovered with less than 6 months lead time often cannot be resolved before they impact operations.

Traffic Analysis

Traffic Measurement

Accurate capacity planning requires accurate traffic measurement. Key metrics:

  • Average utilization: Average bandwidth utilization over a measurement period (typically 5-minute intervals)
  • Peak utilization: Maximum utilization observed during the measurement period
  • 95th percentile utilization: The utilization level exceeded only 5% of the time — the standard metric for capacity planning
  • Traffic direction: Ratio of east-west to north-south traffic
  • Traffic patterns: Time-of-day and day-of-week patterns

Traffic Sources

Collect traffic data from: SNMP polling of switch interfaces (5-minute intervals minimum), NetFlow/IPFIX for traffic flow analysis, and application performance monitoring for application-level traffic patterns.

Traffic Growth Trends

Analyze historical traffic trends to project future growth. Data center east-west traffic has grown 30–50% annually in recent years, driven by microservices, distributed databases, and AI workloads. Apply growth rates to current measurements to project future requirements.

Bandwidth Modeling

Utilization Targets

Target utilization levels for capacity planning:

  • Core/spine links: 60–70% average utilization; upgrade when 95th percentile exceeds 80%
  • Leaf uplinks: 50–60% average utilization; upgrade when 95th percentile exceeds 75%
  • Server access links: 40–50% average utilization; upgrade when 95th percentile exceeds 70%
  • WAN/internet links: 50–60% average utilization; upgrade when 95th percentile exceeds 75%

Oversubscription Modeling

Model oversubscription ratios based on actual traffic patterns. If servers generate 10 Gbps of traffic but only 2 Gbps is destined for other servers (east-west), a 5:1 oversubscription ratio may be acceptable. If 8 Gbps is east-west traffic, a 1.25:1 oversubscription ratio is required.

Burst Capacity

Plan for traffic bursts that exceed average utilization. Batch jobs, backup windows, and application deployments can cause traffic spikes 3–5x average utilization. Ensure sufficient headroom for expected burst patterns.

Port Density Planning

Port density planning ensures that sufficient switch ports are available for current and future server deployments.

Current Inventory

Document current port utilization: total ports, used ports, and available ports by switch and by rack. Identify switches approaching port exhaustion.

Growth Projection

Project future port requirements based on: planned server deployments, storage system additions, and network appliance additions. Include ports for management, out-of-band, and redundant connections.

Port Speed Upgrades

Plan for port speed upgrades as server NIC speeds increase. The industry is transitioning from 25GbE to 100GbE server connections; AI servers require 400GbE. Leaf switches must support the port speeds required by current and future servers.

AI Workload Capacity

AI training workloads require fundamentally different capacity planning than traditional workloads:

All-Reduce Traffic

During all-reduce operations, every GPU communicates with every other GPU simultaneously. A 64-GPU cluster generates 64 × 400 Gb/s = 25.6 Tb/s of all-reduce traffic. The network fabric must support this traffic without congestion.

Non-Blocking Requirement

AI training requires non-blocking (1:1 oversubscription) fabric. Traditional capacity planning with 2:1 or 4:1 oversubscription is insufficient — any congestion causes GPU stalls and reduces training efficiency.

Storage Bandwidth

Training data must be delivered to GPUs at sufficient throughput to keep them fed. Plan storage network bandwidth separately from compute fabric: 5–10 GB/s per 8-GPU server for typical training workloads.

Monitoring & Alerting

Capacity planning is only as good as the monitoring data it is based on. Key monitoring requirements:

  • Interface utilization: 5-minute polling of all switch interfaces via SNMP or streaming telemetry
  • Error counters: Monitor for interface errors, discards, and CRC errors that indicate physical layer problems
  • Queue depth: Monitor switch queue depths for congestion indicators
  • Flow data: NetFlow/IPFIX for traffic flow analysis and anomaly detection

Alerting Thresholds

Set alerts at: 70% utilization (warning), 85% utilization (critical). Alerts should trigger capacity review, not emergency action — capacity planning should prevent reaching critical thresholds.

Upgrade Planning

Upgrade Triggers

Initiate upgrade planning when: 95th percentile utilization exceeds 75% on critical links, port utilization exceeds 80%, or new workloads are planned that will significantly increase traffic.

Upgrade Options

  • Link aggregation: Bond multiple links for higher bandwidth (limited by switch port availability)
  • Port speed upgrade: Replace 10GbE with 25GbE, or 100GbE with 400GbE
  • Additional spine switches: Add spine switches to increase bandwidth between leaf switches
  • Additional leaf switches: Add leaf switches to add server capacity
  • Full fabric replacement: Replace the entire fabric for major architecture changes

Upgrade Execution

Network upgrades in live environments require careful planning: maintenance windows, rollback procedures, and staged rollouts. Test upgrades in non-production environments before applying to production. Document all changes and verify performance after each change.