Evaluation Criteria
Performance
- Forwarding rate (packets per second) at line rate
- Latency (cut-through vs. store-and-forward)
- Buffer size (critical for bursty traffic patterns)
- ECMP support and hash distribution quality
- Throughput with all features enabled (ACLs, QoS, monitoring)
Scalability
- Maximum port density per switch
- Maximum fabric size (number of switches)
- Route table capacity (IPv4, IPv6, EVPN MAC/IP)
- VXLAN tunnel capacity
Software and Automation
- Programmability (Python, Ansible, Terraform support)
- Streaming telemetry (gNMI/gRPC)
- Zero-touch provisioning
- Integration with orchestration platforms (Kubernetes, VMware)
- Software update frequency and stability
Support and Ecosystem
- Support SLAs (response time, hardware replacement)
- Software support lifecycle
- Partner ecosystem and professional services availability
- Vendor financial stability
Switch Evaluation
Leaf Switch Evaluation
Key leaf switch criteria: port density (48x 25GbE + 8x 100GbE is standard; 32x 100GbE + 8x 400GbE for high-density), forwarding rate at line rate, buffer size for bursty workloads, and VXLAN/EVPN support.
Spine Switch Evaluation
Key spine switch criteria: port count (32x or 64x 400GbE), non-blocking architecture, ECMP hash quality, and BGP route table capacity.
Proof of Concept
Test switches in your environment before committing. Key PoC tests: forwarding rate at line rate with all features enabled, ECMP hash distribution quality, failover time, and automation integration with your tooling.
AI Networking Evaluation
AI networking evaluation requires different criteria than traditional enterprise networking:
- All-reduce performance: Test with actual AI training workloads (NCCL all-reduce benchmarks), not synthetic traffic generators
- Non-blocking architecture: Verify that the fabric is truly non-blocking at the required scale
- RDMA support: For RoCEv2, verify PFC and ECN configuration and test under congestion
- NVIDIA SHARP support: For InfiniBand, verify SHARP support for in-network all-reduce
- Scaling efficiency: Measure GPU utilization during distributed training — network bottlenecks reduce GPU utilization
Managed Network Services
Managed network services provide ongoing network operations support. Evaluation criteria:
- Scope of management (monitoring only, or active management and changes?)
- Response time SLAs for incidents and changes
- NOC staffing and expertise
- Tooling and visibility (do you retain access to your network?)
- Change management process
- Reference customers with similar environments
Vendor Comparison
| Vendor | Strengths | Considerations |
|---|---|---|
| Cisco Nexus | Largest installed base; ACI SDN; strong support | Higher cost; complex licensing; ACI learning curve |
| Arista EOS | Programmable; excellent automation; strong in finance | Smaller installed base; fewer professional services partners |
| Juniper QFX | Apstra intent-based; strong in service providers | Smaller data center market share; fewer integrations |
| NVIDIA Spectrum | Best for AI (Spectrum-X); InfiniBand leadership | Newer in enterprise Ethernet; smaller ecosystem |
| White-box (SONiC) | Lowest cost; open-source NOS; vendor independence | Higher operational complexity; limited vendor support |
Common Procurement Mistakes
- Evaluating on spec sheets alone: Lab performance often differs significantly from production performance — test in your environment
- Ignoring software licensing costs: Network OS licenses, feature licenses, and support contracts add significantly to hardware costs
- Underestimating operational complexity: Some platforms require significant expertise to operate effectively
- Not planning for AI workloads: Traditional enterprise network criteria are insufficient for AI infrastructure
- Selecting on price alone: The cheapest network is rarely the best value when performance, reliability, and operational costs are considered
- Ignoring vendor stability: Network infrastructure has a 5–7 year lifecycle — vendor financial stability matters