Skip to main content
DCS Global

GPU Cluster Infrastructure Guides

Quick Reference

GPU Clusters — Quick Reference

Definitions, specifications, and decision frameworks for GPU cluster design and deployment.

Definition: GPU Cluster

A GPU cluster is a group of GPU-accelerated servers interconnected with high-bandwidth, low-latency networking (InfiniBand or RoCE) to function as a single computational unit for AI training, HPC, and large-scale inference workloads. GPU clusters enable distributed training of models too large to fit in a single server's GPU memory, and provide the aggregate compute required for production AI at scale.

  • ▸Smallest practical cluster: 8 GPUs (1 server) — suitable for models up to ~70B parameters.
  • ▸Medium cluster: 64–256 GPUs — suitable for training 70B–405B parameter models.
  • ▸Large cluster: 512–4,096+ GPUs — required for frontier model training and hyperscale inference.
  • ▸Interconnect: InfiniBand NDR (400 Gbps) or 400G Ethernet with RoCE v2 for inter-node communication.

GPU Cluster Size Guide

Representative values for NVIDIA H100 SXM5. Actual requirements vary by model architecture and training configuration.
Cluster sizeGPU countGPU memorySuitable forApprox. hardware cost
Single server8 GPUs640 GBModels up to 70B params; inference$2.5M–$4M
Small cluster32–64 GPUs2.5–5 TB70B–405B param training$10M–$20M
Medium cluster128–256 GPUs10–20 TBFrontier model fine-tuning$40M–$80M
Large cluster512–1,024 GPUs40–80 TBFrontier model pre-training$160M–$320M

GPU Cluster Design Decisions

  • GPU selection: NVIDIA H100 SXM5 (highest training throughput), H100 PCIe (lower cost, lower bandwidth), H200 (larger HBM3e memory for larger models), AMD MI300X (vendor diversification).
  • Interconnect: InfiniBand NDR (400 Gbps) for maximum all-reduce performance; 400G Ethernet with RoCE v2 for operational simplicity and lower cost.
  • Storage: parallel file systems (WEKA, GPFS, Lustre) for training data; local NVMe for checkpoints; object storage (S3-compatible) for model artifacts.
  • Power: 10.2 kW per H100 SXM5 server; 82 kW per 8-GPU cluster; 110–120 kW total per cluster including networking and cooling overhead.
  • Cooling: direct liquid cooling (DLC) required above 15 kW/rack; rear-door heat exchangers for 10–20 kW/rack; immersion for 30+ kW/rack.
  • Management: Kubernetes with GPU operator for job scheduling; Slurm for HPC-style workloads; DCGM for GPU health monitoring.
6 Articles

GPU Cluster Infrastructure

Cluster architecture, InfiniBand, distributed training, and job scheduling. Complete technical guides for designing and operating GPU clusters from 8 to 1,000+ GPUs.

Reference

Key Concepts

Fabric Design

InfiniBand and Ethernet fabric topologies for GPU clusters — fat-tree, rail-optimized, and the trade-offs between bandwidth, latency, and cost.

InfiniBand

RDMA-capable interconnect for GPU clusters — HDR (200Gb/s), NDR (400Gb/s), and the configuration required for optimal all-reduce performance.

Distributed Training

Data, tensor, and pipeline parallelism strategies — and the cluster infrastructure required to support each training parallelism approach.

Job Scheduling

Slurm, Kubernetes, and Run:ai for GPU cluster scheduling — queue management, gang scheduling, and utilization optimization strategies.

Ready to Build Your AI Infrastructure?

Our certified engineers design and deploy enterprise AI infrastructure — from single GPU servers to 1,000+ GPU clusters.