Basic Questions

Q: What is the difference between a data center and a server room?

A server room is a small space within a building dedicated to IT equipment — typically a converted office or utility room with basic cooling and power. A data center is a purpose-built facility with engineered power, cooling, physical security, and fire suppression systems designed for continuous, mission-critical operation. Data centers are designed to Uptime Institute Tier standards; server rooms typically are not.

Q: What is colocation?

Colocation (colo) is a service where a data center provider rents space, power, and connectivity to customers who own their own IT equipment. The provider manages the facility infrastructure (power, cooling, physical security, connectivity); the customer manages their own servers, storage, and networking equipment within their leased space.

Q: How much does a data center cost to build?

Data center construction costs vary widely by size, tier, and location. Rough estimates: $8–$15 million per MW for Tier III construction in the US (2024 pricing). A 1 MW data center costs $8–$15 million to build, plus land, utility infrastructure, and equipment. Operating costs add $1–$3 million per MW per year. These costs make building your own data center economically viable only at significant scale (10+ MW).

Q: What is PUE and what is a good PUE?

PUE (Power Usage Effectiveness) = Total Facility Power ÷ IT Equipment Power. A PUE of 1.0 is theoretical perfection. Industry average is approximately 1.5. A PUE below 1.4 is good; below 1.2 is excellent. Google's data centers average 1.10. PUE is the primary efficiency metric for data centers and directly affects operating cost.

Q: What is DCIM?

Data Center Infrastructure Management — software that provides real-time visibility and management of data center infrastructure including power, cooling, space, and connectivity. DCIM enables proactive capacity management, energy optimization, and change management. Leading DCIM platforms include Nlyte, Sunbird, and Vertiv.

Power Questions

Q: How long can a data center run on UPS power?

UPS runtime depends on battery capacity and load. Typical UPS systems provide 5–15 minutes of runtime at full load — sufficient to bridge the gap until generators start (10–15 seconds). Some facilities use extended battery modules for 30–60 minutes of runtime. Flywheel UPS systems provide only 10–30 seconds but with faster response and no battery degradation.

Q: What happens when a data center loses utility power?

In a properly designed Tier III/IV facility: UPS systems provide instantaneous power (zero transfer time for double-conversion UPS), generators start within 10–15 seconds, automatic transfer switches connect generators to the facility bus, and UPS batteries bridge the gap. The entire sequence takes 10–30 seconds; IT equipment experiences no interruption.

Q: What is N+1 redundancy?

N+1 means one additional component beyond the minimum required. If 3 UPS modules are needed to carry the load (N=3), 4 are installed (N+1=4). Any single module can fail without affecting operations. N+1 is the minimum redundancy for Tier III concurrent maintainability.

Q: What is 2N redundancy?

2N means two complete, independent systems — each capable of carrying the full load. Required for Tier IV fault tolerance. Significantly higher cost than N+1 but provides protection against any single failure, including failures in the distribution path (not just component failures).

Cooling Questions

Q: What temperature should a data center be?

ASHRAE recommends server inlet temperatures of 18–27°C (64–80°F) for A1 class equipment. Modern equipment (ASHRAE A3/A4) can operate at inlet temperatures up to 40–45°C, enabling higher cooling water temperatures and improved efficiency. Cold aisle temperatures are typically set at 18–22°C in most facilities.

Q: What is hot/cold aisle containment?

Hot/cold aisle containment separates hot server exhaust air from cold supply air to prevent mixing. Cold aisle containment encloses the cold aisle; hot aisle containment encloses the hot aisle. Containment improves cooling efficiency by 20–40% and is one of the most cost-effective cooling improvements available.

Q: When do I need liquid cooling?

Liquid cooling becomes necessary above approximately 20–30 kW per rack for air-cooled facilities. For AI infrastructure (H100/H200 servers at 8–10 kW each), liquid cooling is strongly recommended above 40 kW per rack. NVIDIA GB200 NVL72 systems at 120 kW per rack require direct liquid cooling — air cooling is not supported.

Q: What is free cooling?

Free cooling (economizer mode) uses outdoor air or water to cool the data center when ambient temperatures are low enough, bypassing or supplementing mechanical refrigeration. Can dramatically reduce cooling energy consumption in cooler climates — some facilities achieve 70–80% free cooling hours annually. Requires careful humidity and contamination management.

Tier Questions

Q: What Tier data center do I need?

For most enterprise production workloads: Tier III (99.982% availability, 1.6 hours downtime/year). Tier IV is appropriate only when the cost of 26 minutes of annual downtime exceeds the 25–50% premium over Tier III. Tier I and II are appropriate only for non-critical applications where planned maintenance windows are acceptable.

Q: Is Tier IV always better than Tier III?

Not necessarily. Tier IV provides fault tolerance (any single failure cannot affect IT operations) but at significantly higher cost and operational complexity. For most workloads, Tier III's concurrent maintainability (maintenance without downtime) provides sufficient availability. Tier IV's additional protection against unplanned failures is only worth the premium for the most critical applications.

Q: Can a data center lose its Tier certification?

Yes. Uptime Institute's Operational Sustainability certification requires ongoing compliance with operational standards. Facilities that fail to maintain their infrastructure, staffing, or procedures can lose certification. Design and Constructed Facility certifications are permanent once issued but do not guarantee operational performance.

Colocation Questions

Q: How much does colocation cost?

Colocation pricing varies by market, provider, and commitment level. Typical retail colocation: $100–$300 per cabinet per month for basic power (2–3 kW); $500–$2,000 per cabinet per month for high-density (10–20 kW). Wholesale colocation (250+ kW): $80–$150 per kW per month. Power costs are often separate from space costs.

Q: What is a cross-connect?

A cross-connect is a physical cable connection between two customers within the same colocation facility, or between a customer and a carrier. Cross-connects enable direct, low-latency connectivity without traversing the public internet. Essential for financial services, content delivery, and any application requiring predictable, low-latency connectivity.

Q: What is a meet-me room (MMR)?

A meet-me room is a neutral space within a colocation facility where multiple carriers and customers can interconnect. The MMR is the hub of the facility's connectivity ecosystem — the more carriers present in the MMR, the more connectivity options available to customers.

AI Infrastructure Questions

Q: Can I deploy AI servers in a standard colocation facility?

It depends on the facility's power density capability. Standard colocation facilities support 5–10 kW per cabinet — insufficient for AI servers (8–10 kW per server, 8 servers per rack = 64–80 kW per rack). You need a facility that supports high-density deployments (20+ kW per cabinet) and ideally offers liquid cooling. Verify power density capability before signing a contract.

Q: How much power does an AI training cluster consume?

A 64-GPU H100 cluster (8 servers × 8 GPUs) consumes approximately 80 kW of IT load. With cooling and overhead, total facility power is 100–120 kW. A 512-GPU cluster consumes approximately 640 kW of IT load, requiring 800 kW–1 MW of facility power capacity. Plan facility power requirements before procuring AI hardware.

Q: What is the difference between training and inference infrastructure?

Training infrastructure (H100/H200 clusters with InfiniBand) is optimized for throughput — completing training runs as fast as possible. Inference infrastructure is optimized for latency and cost-per-request — serving model predictions to users or applications. Training clusters are typically batch workloads; inference is continuous, latency-sensitive serving. They have different hardware, networking, and operational requirements.

Operations Questions

Q: What causes most data center outages?

Human error is the leading cause of data center outages — studies consistently show 60–80% of outages are caused by mistakes during maintenance, configuration changes, or emergency response. Equipment failure, utility power outages, and cooling failures account for most of the remainder. This is why change management procedures and staff training are as important as infrastructure redundancy.

Q: How often should generators be tested?

Monthly no-load tests verify starting capability. Quarterly or annual load tests (under actual load) verify performance under real conditions. Generators that are only tested under no-load conditions have significantly higher failure rates during actual outages. Annual full-load tests are the minimum for mission-critical facilities.

Q: What is a hot site vs. cold site for disaster recovery?

A hot site is a fully equipped, operational data center that can take over operations immediately following a disaster at the primary site. A cold site is a facility with power and connectivity but no IT equipment — equipment must be procured and installed after a disaster. Hot sites provide near-zero RTO (Recovery Time Objective) but cost significantly more than cold sites.