A–C

All-Flash Array (AFA)

A storage array that uses flash (SSD) storage exclusively, with no spinning disk. Delivers sub-millisecond latency and high IOPS. Has replaced hybrid arrays as the standard for performance-sensitive workloads. Leading vendors: Pure Storage, NetApp AFF, Dell PowerStore.

Block Storage

Storage that presents raw storage volumes to servers, which format them with a file system. Provides the lowest latency and highest IOPS of any storage type. Used for databases, virtual machines, and high-performance applications. Protocols: Fibre Channel, iSCSI, NVMe-oF.

Ceph

An open-source distributed storage platform providing block (RBD), file (CephFS), and object (RADOS Gateway/S3) storage from a single cluster. Highly scalable; complex to manage. Popular for cloud infrastructure and large-scale deployments.

Compression

A data reduction technique that reduces the size of data by encoding it more efficiently. Lossless compression preserves all data; lossy compression (not used for enterprise storage) discards some data. Compression ratios vary by data type: text (3:1 to 5:1), databases (2:1 to 4:1), already-compressed data (1:1).

D–F

Deduplication

A data reduction technique that eliminates duplicate data blocks, storing only unique blocks. Inline deduplication processes data as it is written; post-process deduplication runs as a background job. Effective for virtual machine storage (3:1 to 5:1) and backup data (5:1 to 20:1). Minimal benefit for AI training data.

Erasure Coding

A data protection method that splits data into fragments and adds parity fragments, enabling recovery from multiple drive or node failures. More storage-efficient than replication: 8+4 erasure coding uses 1.5x the raw data size vs. 3x for 3-way replication. Used in object storage and some all-flash arrays.

Fibre Channel (FC)

A high-speed storage networking protocol designed for storage traffic. Provides dedicated storage network, hardware-enforced lossless delivery, and sub-millisecond latency. FC speeds: 16GFC, 32GFC, 64GFC. Requires dedicated FC infrastructure (HBAs, FC switches). The gold standard for mission-critical storage.

Flash Storage

Non-volatile storage using NAND flash memory. Much faster than spinning disk: sub-millisecond latency vs. 5–10ms for HDD. Types: SLC (highest performance, lowest density), MLC, TLC, QLC (highest density, lowest cost). Enterprise SSDs use MLC or TLC for the best balance of performance and endurance.

G–I

GPFS (IBM Spectrum Scale)

IBM's enterprise parallel file system. Supports petabyte-scale deployments with high throughput. Strong data management, tiering, and governance capabilities. Common in financial services, research, and AI deployments requiring enterprise support.

HBA (Host Bus Adapter)

A hardware adapter that connects a server to a storage network. FC HBAs connect to Fibre Channel SANs; iSCSI HBAs (or standard NICs with iSCSI initiator software) connect to iSCSI storage; NVMe HBAs connect to NVMe-oF storage.

iSCSI

Internet Small Computer Systems Interface — a storage protocol that runs SCSI commands over TCP/IP networks. Enables block storage access over standard Ethernet. Lower cost than Fibre Channel; higher latency (200–500 μs vs. 50–100 μs for FC). Widely used for general enterprise storage.

J–N

LUN (Logical Unit Number)

A logical storage volume presented to a server by a storage array. Servers see LUNs as disk devices and format them with a file system. LUN masking controls which servers can access which LUNs. The fundamental unit of block storage allocation.

Lustre

An open-source parallel file system widely used in HPC and AI. Provides very high aggregate throughput by distributing data across multiple storage servers (OSTs — Object Storage Targets). Complex to manage; common in national laboratories and large research institutions.

MinIO

An open-source, high-performance S3-compatible object storage platform. Runs on commodity hardware; delivers very high throughput (100+ GB/s). Popular for AI dataset management and analytics. Available as open source or enterprise edition with commercial support.

MPIO (Multipath I/O)

Software that manages multiple physical paths between a server and storage, providing both redundancy and load balancing. If one path fails, MPIO automatically routes I/O over the remaining paths. Required for high-availability storage deployments.

NAS (Network Attached Storage)

A storage device or system that provides shared file access over a network using NFS (Linux/Unix) or SMB/CIFS (Windows). NAS is optimized for shared access from multiple clients simultaneously. Not suitable for high-performance block workloads.

NVMe (Non-Volatile Memory Express)

A storage protocol designed for flash memory, connecting storage devices directly to the CPU via PCIe. Provides much lower latency (20–100 μs) and higher IOPS (1,000,000+) than SATA or SAS. The standard for high-performance flash storage in modern servers and all-flash arrays.

O–R

Object Storage

A storage paradigm that stores data as objects with unique identifiers, accessed via HTTP/S3 API. Provides virtually unlimited scalability, built-in redundancy, and low cost per GB. The standard for AI datasets, backup, and unstructured data at scale.

RAID (Redundant Array of Independent Disks)

A data storage virtualization technology that combines multiple drives for redundancy and/or performance. Common RAID levels: RAID 1 (mirroring), RAID 5 (distributed parity), RAID 6 (dual parity), RAID 10 (mirrored stripes). Modern all-flash arrays use RAID 5 or RAID 6 equivalents for data protection.

Replication

Copying data to a secondary location for disaster recovery or high availability. Synchronous replication writes to both locations before acknowledging the write (zero RPO, higher latency). Asynchronous replication writes to the primary first, then replicates (some data loss risk, lower latency impact).

S–Z

SAN (Storage Area Network)

A dedicated high-speed network that connects servers to storage systems. SANs use Fibre Channel or iSCSI protocols to provide block storage access. SANs are separate from the LAN (Local Area Network) to ensure storage traffic does not compete with application traffic.

Snapshot

A point-in-time copy of a storage volume or file system. Snapshots are space-efficient (only changed blocks are stored separately) and near-instantaneous. Used for backup, testing, and recovery. Snapshots are not a substitute for backup — they are stored on the same system as the source data.

Thin Provisioning

Allocating storage capacity on demand rather than upfront. A 10 TB thin-provisioned volume only consumes physical storage as data is written. Improves storage utilization but requires monitoring to prevent over-commitment (allocating more thin-provisioned capacity than physical storage available).

WEKA

A modern parallel file system designed for AI and analytics workloads. Runs on NVMe SSDs; delivers very high throughput with low latency. Cloud-native architecture supports on-premises, cloud, and hybrid deployments. Popular for AI training clusters requiring maximum storage performance.