Compute Requirements

GPU selection is the most consequential infrastructure decision for enterprise AI. The right GPU depends on workload type, model size, throughput requirements, and budget constraints.

Training GPUs

For large-scale model training and fine-tuning, NVIDIA H100 SXM5 (80GB HBM3) remains the production standard as of 2026. Key specifications:

  • 989 TFLOPS FP16 / 3,958 TOPS INT8
  • 80 GB HBM3 with 3.35 TB/s memory bandwidth
  • NVLink 4.0 at 900 GB/s bidirectional (within DGX H100)
  • 700W TDP (SXM5 form factor)

The NVIDIA H200 SXM5 improves on H100 with 141 GB HBM3e memory and 4.8 TB/s bandwidth — critical for models that exceed H100 memory capacity. For new deployments, H200 is preferred for training workloads where memory capacity is a constraint.

NVIDIA Blackwell (B200/GB200) represents the next generation with 2x–4x performance improvements over H100. GB200 NVL72 systems (72 GPUs in a single rack) are designed for the largest training runs and inference deployments.

Inference GPUs

Inference workloads prioritize throughput and cost-per-token over raw training performance. NVIDIA L40S (48GB GDDR6) offers better inference economics than H100 for many serving workloads. For latency-sensitive applications, A100 80GB remains widely deployed. The H100 NVL (94GB) is optimized for large-model inference.

CPU Infrastructure

AI servers require high-core-count CPUs for data preprocessing, orchestration, and model serving overhead. AMD EPYC (Genoa/Bergamo) and Intel Xeon (Sapphire Rapids/Emerald Rapids) are the primary options. For GPU servers, CPU selection is secondary to GPU and interconnect specifications.

Networking Requirements

Network fabric design is the most frequently underestimated component of AI infrastructure. For distributed training, network bandwidth directly limits scaling efficiency — a poorly designed network can reduce GPU utilization to 40–60% of theoretical maximum.

Intra-Node Interconnect

Within a single server, GPUs communicate via NVLink. NVIDIA DGX H100 provides NVLink 4.0 at 900 GB/s bidirectional bandwidth between all 8 GPUs. This is non-configurable — it comes with the DGX platform.

Inter-Node Interconnect

Between servers, two options dominate:

  • InfiniBand NDR (400 Gb/s): The gold standard for AI training. NVIDIA Quantum-2 switches provide non-blocking fat-tree topologies. Required for the largest training runs. Higher cost but superior performance for all-reduce operations.
  • RoCEv2 over 400GbE: RDMA over Converged Ethernet. Lower cost than InfiniBand, uses standard Ethernet infrastructure. Performance gap vs. InfiniBand has narrowed significantly. Viable for many enterprise training workloads.

Storage Network

Storage fabric must deliver sufficient bandwidth to keep GPUs fed during training. A 64-GPU cluster training a large language model may require 400–800 GB/s aggregate storage throughput. This typically requires a dedicated high-speed storage network (100GbE or 200GbE per storage node) separate from the compute fabric.

Management Network

A separate out-of-band management network (1GbE or 10GbE) for BMC/IPMI access, DCIM integration, and infrastructure management. This network must be isolated from the compute and storage fabrics.

Storage Requirements

AI storage requirements differ from traditional enterprise storage in three key ways: throughput (not IOPS) is the primary metric, access patterns are highly sequential during training, and checkpoint storage requires high write bandwidth.

Training Data Storage

Training datasets for large models can range from tens of terabytes to petabytes. Parallel file systems (GPFS/IBM Spectrum Scale, Lustre, WEKA) provide the throughput required for GPU-dense clusters. NFS is generally insufficient for large training clusters.

Minimum throughput guidelines:

  • 8-GPU node: 50–100 GB/s aggregate read throughput
  • 32-GPU cluster: 200–400 GB/s aggregate read throughput
  • 128-GPU cluster: 800 GB/s – 1.6 TB/s aggregate read throughput

Checkpoint Storage

Model checkpoints during training must be written quickly to minimize GPU idle time. A 70B parameter model checkpoint is approximately 140 GB (FP16). With checkpointing every 30 minutes, this requires sustained write throughput of 5–10 GB/s minimum.

Model Repository

Trained models, model versions, and serving artifacts require reliable, versioned storage. Object storage (MinIO, Ceph, or cloud-compatible S3-compatible storage) is standard for model repositories.

NVMe Local Storage

Each AI server should have local NVMe storage for dataset caching, temporary files, and OS. Minimum 4x NVMe drives per server (8–16 TB total) for large training workloads.

Power & Cooling Requirements

AI infrastructure has dramatically higher power density than traditional IT. Planning for power and cooling is critical — undersizing either will limit cluster performance or require expensive retrofits.

Power Requirements

  • NVIDIA DGX H100: 10.2 kW per system
  • NVIDIA DGX H200: 10.2 kW per system
  • NVIDIA GB200 NVL72 rack: 120 kW per rack
  • Typical AI server (8x H100 HGX): 8–10 kW per server

A 10-rack AI cluster with 8 servers per rack at 10 kW each requires 800 kW of IT load capacity, plus overhead for networking, storage, and cooling infrastructure. Total facility power requirement: 1.2–1.6 MW for a 10-rack AI cluster.

Cooling Requirements

Traditional air cooling (CRAC/CRAH) is insufficient for AI-dense deployments above 15–20 kW per rack. Options:

  • Rear-door heat exchangers: Capture heat at the rack level. Effective up to 30–40 kW per rack. Lowest infrastructure change required.
  • Direct liquid cooling (DLC): Cold plates on CPUs and GPUs. Required for GB200 NVL72 systems. Supports 50–120 kW per rack.
  • Immersion cooling: Servers submerged in dielectric fluid. Highest density (100+ kW per tank), lowest PUE, but highest infrastructure cost and operational complexity.

NVIDIA GB200 NVL72 systems require direct liquid cooling — air cooling is not supported. Plan cooling infrastructure before procuring next-generation AI hardware.

Inference Infrastructure

Production inference has different requirements than training. Key differences:

  • Latency SLAs (50–200ms for interactive applications) vs. throughput optimization for training
  • Continuous availability requirements (99.9%+ uptime) vs. batch training jobs
  • Variable load patterns requiring auto-scaling capabilities
  • Cost-per-token optimization as the primary economic metric

Inference serving stacks (vLLM, TensorRT-LLM, NVIDIA Triton) use techniques like continuous batching, KV cache management, and tensor parallelism to maximize GPU utilization. Proper serving stack configuration can improve throughput by 5–10x compared to naive inference.

For large language model inference, memory capacity is often the binding constraint. A 70B parameter model in FP16 requires 140 GB of GPU memory — requiring either multi-GPU serving (tensor parallelism) or quantization (INT8/INT4) to fit on available hardware.

Sizing Your Cluster

Cluster sizing depends on model size, training duration targets, and inference throughput requirements. A simplified sizing framework:

Training Cluster Sizing

For fine-tuning a 70B parameter model: 8x H100 80GB GPUs minimum (using LoRA/QLoRA), 16–32x H100 for full fine-tuning. For pre-training a 7B parameter model from scratch: 64–128x H100 for reasonable training duration (weeks, not months).

Inference Cluster Sizing

Calculate required throughput (tokens per second), determine per-GPU throughput for your model and serving stack, add redundancy (N+1 minimum), and size accordingly. A single H100 80GB can serve approximately 1,000–3,000 tokens/second for a 7B parameter model with optimized serving.

Procurement Considerations

AI infrastructure procurement requires careful attention to lead times, vendor relationships, and total system integration:

  • GPU lead times: H100/H200 systems have historically had 6–12 month lead times. Plan procurement 9–12 months ahead of deployment target.
  • System integration: DGX systems come pre-integrated; HGX-based servers require integration with chassis, networking, and storage. Factor integration time into project schedules.
  • Software stack: NVIDIA AI Enterprise software licensing, CUDA, cuDNN, and framework licenses add to total cost. Include in TCO modeling.
  • Support contracts: Mission-critical AI infrastructure requires 4-hour or next-business-day hardware support. Factor support costs into 3-year TCO.
  • Warranty and refresh cycles: GPU hardware typically has a 3-year warranty. Plan for refresh cycles in long-term infrastructure roadmaps.