Basic Questions
Q: What is the difference between block, file, and object storage?
Block storage presents raw storage volumes to servers (like a disk). File storage provides shared file systems (NAS). Object storage stores data as objects with unique IDs, accessed via HTTP/S3 API. Block storage provides the lowest latency for databases and VMs. File storage is for shared access. Object storage is for unstructured data at scale.
Q: Should I use all-flash or hybrid storage?
For most new deployments, all-flash is the right choice. All-flash prices have declined dramatically — for performance-sensitive workloads, all-flash is now cost-competitive with hybrid when TCO is considered. Hybrid arrays (flash + spinning disk) are only appropriate for workloads with very low performance requirements and very high capacity requirements.
Q: What is software-defined storage?
Software-defined storage (SDS) decouples storage software from proprietary hardware, enabling deployment on commodity servers. Examples: Ceph, VMware vSAN, Nutanix. SDS provides flexibility and lower hardware cost but requires more operational expertise than purpose-built storage arrays.
Performance Questions
Q: What storage performance do I need for my database?
Database storage requirements depend on the database type and workload. OLTP databases (Oracle, SQL Server) require high IOPS and low latency — all-flash block storage with sub-millisecond latency. OLAP/analytics databases require high throughput — all-flash or NVMe-oF. As a starting point: 10,000–100,000 IOPS and sub-1ms latency for OLTP; 10–100 GB/s throughput for analytics.
Q: What is the difference between IOPS and throughput?
IOPS (Input/Output Operations Per Second) measures the number of read/write operations per second — relevant for random access workloads (databases, VMs). Throughput (MB/s or GB/s) measures the amount of data transferred per second — relevant for sequential access workloads (AI training, backup, video). Most workloads require both, but one is typically the binding constraint.
Q: Why does storage performance degrade at high utilization?
Storage performance degrades above 70–80% utilization because: garbage collection (on flash storage) competes with I/O, write amplification increases, and cache effectiveness decreases. Plan storage capacity to maintain utilization below 75% for consistent performance.
AI Storage Questions
Q: What storage do I need for AI training?
AI training requires parallel file systems (GPFS, Lustre, WEKA) for the training data path — traditional NAS is insufficient for GPU-dense clusters. A 64-GPU cluster may require 400–800 GB/s aggregate throughput. Object storage (S3-compatible) is used for dataset management. Local NVMe SSDs are used for checkpoint storage.
Q: Can I use NFS for AI training?
Standard NFS is insufficient for GPU-dense AI training clusters. NFS typically delivers 10–50 GB/s aggregate throughput — far below the 400+ GB/s required for large clusters. NFSv4.1 with pNFS improves throughput but is still limited compared to parallel file systems. Use parallel file systems (WEKA, GPFS, Lustre) for AI training.
Q: How much storage do I need for AI?
AI storage requirements depend on dataset size, model size, and retention requirements. A rough guide: training datasets (10–100 TB for typical enterprise AI), model checkpoints (1–10 TB per training run), experiment results (1–10 TB per project), and model artifacts (100 GB–1 TB per model). Plan for 50–100% annual growth in AI storage requirements.
Protocol Questions
Q: Should I use Fibre Channel or iSCSI?
Fibre Channel provides higher performance and reliability but at higher cost. iSCSI runs over standard Ethernet at lower cost with good performance. For mission-critical databases and applications where latency is critical, FC is preferred. For general enterprise storage, iSCSI over 25GbE or 100GbE provides good performance at lower cost. For new deployments, consider NVMe/TCP as a modern alternative to both.
Q: What is NVMe-oF and when should I use it?
NVMe-oF (NVMe over Fabrics) extends NVMe over the network, providing near-local NVMe performance for shared storage. Use NVMe-oF when: latency below 200 microseconds is required, throughput above 10 GB/s per connection is needed, or you are deploying new storage infrastructure and want the best performance. NVMe/TCP is the most accessible option — no specialized hardware required.
Planning Questions
Q: How do I calculate storage capacity requirements?
Calculate: current used capacity + (monthly growth rate × planning horizon in months) + buffer (25–30%). Apply data reduction ratios (deduplication, compression) to get raw capacity requirements. For AI workloads, assume minimal data reduction (1:1 to 1.5:1). For virtual machines, assume 3:1 to 5:1 data reduction on all-flash arrays.
Q: How often should I review storage capacity?
Review storage capacity monthly for fast-growing environments (AI, analytics), quarterly for stable environments. Set automated alerts at 70% utilization (warning) and 80% (critical). Initiate procurement when utilization reaches 75% — allow 3–6 months for procurement and deployment.