OEM Evaluation
Data Center Switching OEM Evaluation
| OEM | Platform | Automation | AI/HPC | Support |
|---|---|---|---|---|
| Cisco | Nexus 9000 | NSO, ACI, NX-API | Strong | Excellent TAC |
| Arista | EOS (7000/7800 series) | eAPI, CloudVision, Ansible | Very strong | Strong TAC |
| NVIDIA | Spectrum (SN series) | NVIDIA Air, Ansible | Purpose-built | Good (AI-focused) |
| Juniper | QFX series | Apstra, PyEZ, Ansible | Strong | Good TAC |
| HPE Aruba | CX series | Fabric Composer, Ansible | Adequate | Good |
Key Specifications
Forwarding capacity (Tbps)
Total switching bandwidth. A 32-port 400 GbE switch has 12.8 Tbps of forwarding capacity. Verify that the switch is non-blocking, forwarding capacity equals the sum of all port speeds.
Buffer depth
The amount of memory available to buffer packets during congestion. Deep buffers are critical for AI training workloads that generate bursty traffic. Shallow-buffer switches drop packets during bursts, degrading AI training performance.
Latency
Cut-through switching provides ~300 nanosecond latency. Store-and-forward switching provides ~1–2 microsecond latency. For AI training, cut-through switching is preferred.
ECMP paths
The number of equal-cost paths the switch can load-balance across. More ECMP paths provide better load distribution in large spine-leaf fabrics.
PFC and ECN support
Required for lossless Ethernet (RoCEv2). Verify that the switch supports Priority Flow Control (PFC) and Explicit Congestion Notification (ECN) on all ports.
AI Networking Requirements
AI training clusters require network infrastructure that is qualitatively different from traditional enterprise networking. The key requirements are: high bandwidth (200–400 Gb/s per node), low latency (sub-microsecond), lossless transport (no packet drops under congestion), and RDMA support (for GPU-to-GPU communication without CPU involvement).
Evaluate AI networking separately from general enterprise networking
RFP Requirements
Non-blocking forwarding capacity
Require documentation that the switch is non-blocking, forwarding capacity equals the sum of all port speeds at line rate.
Buffer depth specification
Require buffer depth specifications for all switch models. Deep-buffer switches (>32 MB) are required for AI workloads.
Automation API documentation
Require documentation of the automation API: REST, gRPC, NETCONF, and sample Ansible playbooks for common operations.
Software licensing model
Require complete software licensing cost over the planned lifecycle, including all features required for the deployment.
Contract Terms
Software update commitment
Require a commitment to software updates for the planned lifecycle, including security patches and feature updates. Switches that stop receiving software updates before end of hardware life create security vulnerabilities.
TAC support quality
Require documentation of TAC support response times and escalation paths. Test TAC quality during the evaluation process, not after purchase.
Hardware replacement
Require next-business-day hardware replacement for all production switches. Verify that the vendor maintains local parts inventory.
Migration assistance
Require the vendor to provide migration assistance for the transition from legacy infrastructure, including configuration migration tools and professional services support.