Power Architecture
Power architecture is the most consequential infrastructure decision in a modernization program. The redundancy topology determines the facility's availability ceiling: no amount of compute, network, or storage redundancy can compensate for a single-path power architecture that creates a facility-wide single point of failure.
Power Redundancy Topologies
| Topology | Description | Availability | Best For |
|---|---|---|---|
| N | Single path, no redundancy | 99.671% (28.8 hrs/yr downtime) | Development, test environments |
| N+1 | One redundant component per system | 99.741% (22.7 hrs/yr) | General enterprise workloads |
| 2N | Fully redundant dual-path power | 99.982% (1.6 hrs/yr) | Mission-critical production |
| 2N+1 | Dual-path plus additional redundancy | 99.995% (26 min/yr) | Tier IV, financial services, healthcare |
UPS selection is the second critical power decision. Double-conversion UPS provides the cleanest power and the best protection against power quality events, but carries higher energy cost than line-interactive designs. Modular UPS systems allow capacity to be added in increments, which is important for facilities that expect workload growth. Battery technology selection: VRLA vs. lithium-ion: affects runtime, maintenance requirements, and total cost of ownership over the UPS lifecycle.
Generator sizing
Cooling Systems
Cooling technology selection is driven by power density. The relationship is direct: higher power density requires more cooling capacity per square foot, and above certain thresholds, air cooling cannot physically remove heat fast enough to maintain safe operating temperatures.
Perimeter Air Cooling (CRAC/CRAH)
Up to 5 kW/rackAdequate for traditional IT workloads. Inefficient for high-density deployments, hot aisle/cold aisle containment required to prevent hot spots.
In-Row Cooling
5–15 kW/rackPlaces cooling units between server rows, reducing the distance hot air must travel. More efficient than perimeter cooling for medium-density deployments.
Rear-Door Heat Exchangers
10–30 kW/rackAttaches to the rear of server racks and captures heat at the source. Requires chilled water infrastructure. Effective for high-density compute.
Direct-to-Chip Liquid Cooling
20–60 kW/rackDelivers coolant directly to CPU and GPU heat spreaders. Required for sustained GPU operation at full TDP. Requires facility-side liquid infrastructure.
Immersion Cooling
100+ kW/rackSubmerges servers in dielectric fluid. Highest density capability, lowest PUE, but requires specialized hardware and significant facility modification.
Network Architecture
Modern data center network architecture has converged on the spine-leaf topology as the standard for enterprise deployments. Spine-leaf provides predictable latency (every server is exactly two hops from every other server), horizontal scalability (adding leaf switches adds capacity without redesigning the spine), and simplified operations (consistent topology reduces configuration complexity).
The three-tier architecture (access, distribution, core) that was standard in enterprise networks through the 2010s is inadequate for modern east-west traffic patterns. Applications that communicate heavily between servers: distributed databases, microservices, AI training clusters, generate east-west traffic that three-tier architectures were not designed to carry efficiently.
Bandwidth planning for AI workloads
Compute Platform
Compute platform selection must account for the full workload mix: not just the dominant workload type. Most enterprise data centers run a combination of CPU-bound workloads (databases, ERP, web applications), memory-intensive workloads (in-memory databases, analytics), and increasingly, GPU-accelerated workloads (AI/ML, scientific computing, rendering).
Compute Platform Selection by Workload Type
| Workload Type | Platform | Key Specifications |
|---|---|---|
| General enterprise (ERP, web, database) | High-core-count CPU servers | 2-socket, 32–64 cores/socket, 512 GB–2 TB RAM |
| In-memory analytics | High-memory CPU servers | 4–8 socket, 6–12 TB RAM, NVMe-backed swap |
| AI training | GPU servers (8x H100/H200) | 8 GPUs, 700W TDP each, NVLink fabric, 200 Gb/s NIC |
| AI inference | GPU or inference accelerator | Lower GPU count, optimized for throughput/latency ratio |
| Edge compute | Ruggedized compact servers | Low power, wide temperature range, remote management |
Storage Design
Storage architecture must match the access patterns of the workloads it serves. The most common storage design mistake is applying a single storage tier to all workloads: resulting in either overspending on high-performance storage for cold data, or underperforming on hot data because the storage tier was sized for cost rather than performance.
Modern storage architecture uses tiering: NVMe all-flash for hot data requiring sub-millisecond latency, SAS/SATA SSD for warm data requiring consistent throughput, and object storage for cold data requiring cost-effective capacity. Automated tiering moves data between tiers based on access frequency, reducing the manual management burden.
DCIM and Operations Infrastructure
Data Center Infrastructure Management (DCIM) software provides the operational visibility that modern facilities require: real-time monitoring of power consumption, cooling performance, environmental conditions, and asset inventory. Without DCIM, capacity planning is guesswork, energy optimization is impossible, and incident response is reactive rather than proactive.
DCIM implementation is not a post-modernization activity, it should be designed into the modernization program from the beginning. Retrofitting DCIM into a facility that was not designed for it is more expensive and less effective than deploying it as part of the modernization program.