The Power Chain
Data center power follows a defined chain from the utility grid to IT equipment. Each stage in the chain must be designed for the required redundancy level, and each stage introduces potential failure points that must be mitigated.
The six-stage power chain:
- Utility feed: Medium-voltage power from the utility grid (typically 13.8 kV or 34.5 kV)
- Transformers: Step down utility voltage to distribution voltage (typically 480V or 208V)
- Switchgear: Distributes power and provides protection and switching capability
- UPS systems: Conditions power and provides backup during outages
- PDUs (Power Distribution Units): Distribute power to rack rows
- Rack PDUs: Distribute power to individual servers within racks
For Tier III and Tier IV facilities, each stage must have redundant paths — A-side and B-side — so that any single component failure or maintenance activity does not interrupt power to IT equipment.
Utility & Switchgear
Utility Feeds
Mission-critical facilities require multiple utility feeds from different substations or different utility paths. A single utility feed — even with redundant UPS and generators — creates a single point of failure at the utility level. Tier III and Tier IV facilities typically have two utility feeds from geographically diverse substations.
Automatic Transfer Switches (ATS)
ATS devices automatically switch between utility feeds or between utility and generator power when a failure is detected. Transfer time is typically 100–300 milliseconds — fast enough for UPS systems to bridge the gap but not instantaneous. Static transfer switches (STS) provide sub-cycle (less than 4ms) transfer for the most sensitive loads.
Switchgear Design
Main-tie-main (MTM) switchgear configurations allow power to be transferred between utility feeds and generator buses without interruption. Proper switchgear design is critical for achieving concurrent maintainability (Tier III) and fault tolerance (Tier IV).
UPS Systems
UPS (Uninterruptible Power Supply) systems serve two functions: providing instantaneous power during utility outages (bridging until generators start) and conditioning power quality (filtering voltage sags, surges, and harmonics).
UPS Topologies
- Double-conversion (online): IT equipment always runs on UPS output — utility power is continuously converted to DC and back to AC. Provides the best power conditioning and zero transfer time. Most common in mission-critical data centers. Efficiency: 92–97%.
- Line-interactive: UPS conditions power and switches to battery only during outages. Transfer time: 2–4ms. Lower cost than double-conversion but less power conditioning. Appropriate for less critical applications.
- Standby (offline): Passes utility power directly to load; switches to battery during outages. Transfer time: 4–8ms. Lowest cost, minimal conditioning. Not appropriate for mission-critical applications.
Battery Technologies
- VRLA (Valve-Regulated Lead-Acid): Traditional UPS battery. Lower upfront cost, well-understood technology. Typical runtime: 5–15 minutes at full load. Requires replacement every 5–7 years.
- Lithium-ion: Higher upfront cost but longer life (10–15 years), smaller footprint, and better performance at high temperatures. Increasingly adopted for new deployments.
- Flywheel: Stores energy as rotational kinetic energy. Very short runtime (10–30 seconds) but extremely fast response and no battery replacement. Used in conjunction with generators for ride-through.
UPS Sizing
UPS systems must be sized for peak IT load plus growth headroom. Common sizing approach: design for 60–70% of UPS capacity at current load, leaving headroom for growth and N+1 redundancy. Oversized UPS systems operate inefficiently — double-conversion UPS efficiency drops significantly below 40% load.
Generator Systems
Diesel generators provide extended backup power when utility power is unavailable. Generator systems must start and reach full load within 10–15 seconds — fast enough for UPS batteries to bridge the gap.
Generator Sizing
Generators must be sized for the full IT load plus mechanical loads (cooling, lighting, fire suppression). Typical sizing: 125% of maximum IT load to account for motor starting loads and future growth. For N+1 redundancy, each generator must be capable of carrying the full facility load independently.
Fuel Systems
On-site diesel fuel storage provides runtime without fuel delivery. Minimum requirements: 24 hours at full load for most enterprise facilities; 72 hours for critical government and financial facilities. Fuel polishing systems prevent fuel degradation during storage. Fuel delivery contracts ensure replenishment during extended outages.
Generator Testing
Generators must be tested regularly under load to ensure reliability. Monthly no-load tests verify starting capability; quarterly or annual load tests verify performance under actual load conditions. Untested generators have significantly higher failure rates during actual outages.
Paralleling Switchgear
Multiple generators can be paralleled to share load and provide redundancy. Paralleling switchgear synchronizes generator output before connecting them to the bus. This allows N+1 generator configurations where any single generator can be taken offline for maintenance.
Power Distribution
PDUs (Power Distribution Units)
Floor-standing PDUs receive power from the UPS and distribute it to rack rows. Typically include circuit breakers, metering, and monitoring capabilities. For Tier III/IV facilities, each rack row is fed from both A-side and B-side PDUs.
Rack PDUs
Mounted within server racks, rack PDUs distribute power to individual servers. Intelligent rack PDUs provide per-outlet metering, remote switching, and environmental monitoring. Critical for capacity management and power optimization.
Busway Systems
Overhead busway systems distribute power along rack rows, with tap-off boxes providing connections to rack PDUs. More flexible than conduit-based distribution for reconfiguration and capacity changes.
Redundancy Architectures
N+1 Redundancy
One additional component beyond the minimum required to carry the load. If N components are required, N+1 components are installed. Any single component can fail without affecting operations. Required for Tier III concurrent maintainability.
2N Redundancy
Two complete, independent systems — each capable of carrying the full load. Required for Tier IV fault tolerance. Significantly higher cost but provides the highest level of protection against both planned and unplanned outages.
Dual-Corded IT Equipment
Servers and storage systems with dual power supplies connected to A-side and B-side power are essential for Tier III/IV facilities. A failure in either the A-side or B-side power path does not affect IT operations because the other path continues to supply power.
AI Infrastructure Power Requirements
AI servers (8x H100 GPU configurations) draw 8–10 kW per server — 5–10x the density of traditional IT equipment. This creates specific power infrastructure challenges:
- Traditional 20A circuits (2.4 kW) are insufficient — AI servers require 30A or 60A circuits
- Rack-level power requirements of 20–120 kW require high-density PDUs and busway systems
- NVIDIA GB200 NVL72 systems require 120 kW per rack — beyond the capability of most existing power infrastructure
- Power factor correction is critical — AI servers can have poor power factor, increasing apparent power requirements
- Harmonic distortion from switching power supplies requires careful UPS and transformer sizing
Organizations deploying AI infrastructure in existing data centers must conduct a thorough power capacity assessment before procurement. Discovering that the facility cannot support the required power density after hardware delivery is an expensive mistake.
Power Monitoring
Comprehensive power monitoring is essential for capacity management, efficiency optimization, and fault detection:
- Branch circuit monitoring: Per-circuit power consumption for capacity planning and anomaly detection
- Rack-level metering: Total power consumption per rack for density management
- UPS monitoring: Battery health, load percentage, runtime remaining, and power quality metrics
- Generator monitoring: Fuel level, runtime hours, load percentage, and fault conditions
- PUE calculation: Real-time PUE monitoring to track efficiency and identify optimization opportunities
DCIM (Data Center Infrastructure Management) platforms integrate power monitoring with capacity planning, providing visibility into current utilization and projections for future capacity requirements.