AI Storage Requirements

AI training workloads have storage requirements that differ fundamentally from traditional enterprise workloads. The primary metric is throughput (GB/s), not IOPS — training jobs read large sequential blocks of data, not random small blocks.

A 64-GPU H100 cluster training a large language model may require 400–800 GB/s of aggregate storage throughput to keep GPUs fed. Traditional enterprise NAS systems deliver 10–50 GB/s — insufficient for GPU-dense clusters. Parallel file systems are required.

Parallel File Systems

Parallel file systems distribute data across multiple storage nodes and serve it simultaneously from all nodes, achieving aggregate throughput that scales with the number of storage nodes.

  • IBM Spectrum Scale (GPFS): Enterprise-grade parallel file system. Supports petabyte-scale deployments. Strong data management and tiering capabilities. Common in financial services and research.
  • Lustre: Open-source parallel file system widely used in HPC. Very high throughput; complex to manage. Common in national laboratories and large research institutions.
  • WEKA: Modern parallel file system designed for AI workloads. Runs on NVMe SSDs; delivers very high throughput with low latency. Cloud-native architecture.
  • BeeGFS: Open-source parallel file system optimized for performance. Simpler to manage than Lustre. Popular for mid-scale AI deployments.

Object Storage for AI

Object storage (S3-compatible) is the standard for AI dataset management. Datasets are stored as objects, versioned, and accessed via S3 API. Training jobs read datasets from object storage into the parallel file system before training begins.

On-premises S3-compatible object storage: MinIO (open source, high performance), Ceph RADOS Gateway, NetApp StorageGRID. Object storage provides: virtually unlimited scalability, built-in redundancy, rich metadata for dataset management, and low cost per GB.

NVMe and NVMe-oF

NVMe (Non-Volatile Memory Express) is the protocol designed for flash storage, providing much lower latency than SATA or SAS. Local NVMe SSDs in AI servers provide the fastest storage for checkpoints and temporary files.

NVMe-oF (NVMe over Fabrics) extends NVMe over the network — Ethernet (NVMe/TCP) or InfiniBand (NVMe/RDMA). NVMe-oF provides near-local NVMe performance over the network, enabling shared NVMe storage with latency approaching local NVMe.

Three-Tier AI Storage Architecture

The standard AI storage architecture uses three tiers:

  1. Tier 1 — Parallel file system: Active training data. High throughput (200+ GB/s). NVMe-based for lowest latency. Expensive per GB; sized for active datasets only.
  2. Tier 2 — Object storage: Dataset repository, model artifacts, experiment results. Lower throughput; much lower cost per GB. Datasets are staged to Tier 1 before training.
  3. Tier 3 — Archive: Long-term retention of completed experiments, old model versions, and raw data. Tape or cold object storage. Lowest cost per GB.

Checkpoint Storage

Model checkpoints during training must be written quickly to minimize GPU idle time. A 70B parameter model checkpoint is approximately 140 GB (FP16). With checkpointing every 30 minutes, this requires sustained write throughput of 5–10 GB/s minimum.

Local NVMe SSDs provide the fastest checkpoint storage. Each AI server should have 4–8 NVMe drives (8–16 TB total) for checkpoint storage. Checkpoints are then asynchronously copied to the parallel file system or object storage for durability.

Sizing Guide

Cluster SizeRequired ThroughputRecommended Storage
8 GPUs (1 server)50–100 GB/sAll-flash NAS or small WEKA cluster
64 GPUs (8 servers)400–800 GB/sWEKA or GPFS cluster (4–8 nodes)
512 GPUs (64 servers)3–6 TB/sLarge parallel file system (32+ nodes)