A–C

Adversarial Example

A carefully crafted input designed to cause an AI model to produce an incorrect output. Adversarial examples exploit vulnerabilities in model decision boundaries and are used in security testing and adversarial robustness research.

Attention Mechanism

A neural network component that allows models to focus on relevant parts of the input when producing each output token. The foundation of transformer architectures. Self-attention allows each position in a sequence to attend to all other positions.

AutoML

Automated Machine Learning — tools and techniques that automate the process of model selection, hyperparameter tuning, and feature engineering. Reduces the ML expertise required for model development but does not eliminate the need for domain knowledge and governance.

Batch Inference

Processing a large number of inference requests together as a batch, rather than individually. More efficient than real-time inference for non-latency-sensitive workloads. Contrast with online inference (real-time, low-latency serving).

Bias (AI)

Systematic errors in AI model outputs that produce unfair or discriminatory results across demographic groups. Can arise from biased training data, flawed model design, or inappropriate use of model outputs. Distinct from statistical bias (systematic deviation from true values).

CUDA

Compute Unified Device Architecture — NVIDIA's parallel computing platform and programming model that enables GPU-accelerated computation. The dominant programming interface for AI/ML workloads. Most AI frameworks (PyTorch, TensorFlow) use CUDA under the hood.

Concept Drift

A change in the statistical relationship between model inputs and outputs over time, causing model performance to degrade. Distinct from data drift (changes in input distributions). Requires model retraining or recalibration to address.

Continuous Batching

A serving technique for large language models that dynamically adds new requests to in-progress inference batches, rather than waiting for a full batch to form. Dramatically improves GPU utilization and throughput compared to static batching. Used in vLLM and TensorRT-LLM.

D–F

Data Drift

A change in the statistical distribution of model input features over time, relative to the training data distribution. Can cause model performance degradation even when the underlying relationship between inputs and outputs has not changed. Detected using statistical tests (KS test, PSI).

Data Poisoning

An attack in which adversaries inject malicious samples into training data to cause the model to learn incorrect behaviors or create backdoors. A significant security concern for AI systems trained on data from external or untrusted sources.

DGX (NVIDIA)

NVIDIA's purpose-built AI server platform. DGX H100 contains 8x H100 SXM5 GPUs connected via NVLink 4.0 at 900 GB/s bidirectional bandwidth. DGX systems come pre-integrated with NVIDIA's AI software stack. The reference platform for enterprise AI training.

Differential Privacy

A mathematical framework for training AI models while providing provable privacy guarantees — ensuring that the presence or absence of any individual's data cannot be inferred from the model. Adds calibrated noise during training. Used to protect sensitive training data from model inversion attacks.

Embedding

A dense vector representation of data (text, images, audio) in a high-dimensional space where semantically similar items are close together. Embeddings are the foundation of semantic search, recommendation systems, and retrieval-augmented generation.

Explainability

The ability to explain why an AI model produced a particular output in terms understandable to humans. Required for regulatory compliance (GDPR, EU AI Act), model risk management, and debugging. Distinct from interpretability (understanding the model's internal mechanisms).

Feature Store

A centralized repository for storing, managing, and serving ML features — the transformed data representations used as model inputs. Enables feature reuse across models, ensures consistency between training and serving, and provides feature lineage for governance.

Fine-Tuning

Adapting a pre-trained foundation model to a specific task or domain by continuing training on a smaller, task-specific dataset. More cost-effective than training from scratch. Variants include full fine-tuning (all parameters updated), LoRA (low-rank adaptation), and QLoRA (quantized LoRA).

Foundation Model

A large AI model trained on broad data at scale that can be adapted to a wide range of downstream tasks. Examples: GPT-4, Claude, Llama, Mistral. Foundation models are the starting point for most enterprise AI applications through fine-tuning or prompting.

G–I

GPU (Graphics Processing Unit)

A processor designed for parallel computation, originally for graphics rendering but now the dominant hardware for AI/ML workloads. AI-optimized GPUs (NVIDIA H100, H200, AMD MI300X) include specialized tensor cores for matrix multiplication, the core operation in neural network training and inference.

GPTQ

A post-training quantization method for large language models that reduces model precision to INT4 or INT8 with minimal accuracy loss. Enables serving larger models on fewer GPUs. Widely used for deploying 70B+ parameter models on limited GPU memory.

Hallucination

When a large language model generates confident-sounding but factually incorrect or fabricated information. A significant reliability concern for enterprise AI applications. Mitigated through retrieval-augmented generation, output validation, and human-in-the-loop review for high-stakes decisions.

HBM (High Bandwidth Memory)

A type of GPU memory with extremely high bandwidth achieved by stacking memory dies vertically and connecting them with through-silicon vias. NVIDIA H100 uses HBM3 (3.35 TB/s bandwidth); H200 uses HBM3e (4.8 TB/s). HBM bandwidth is often the binding constraint for large model inference.

Hyperparameter

A parameter that controls the training process rather than being learned from data. Examples: learning rate, batch size, number of layers, dropout rate. Hyperparameter tuning (finding optimal values) is a significant part of model development effort.

InfiniBand

A high-speed networking technology used for inter-node communication in AI training clusters. NVIDIA Quantum-2 switches support NDR InfiniBand at 400 Gb/s per port. Provides lower latency and higher bandwidth than Ethernet for all-reduce operations in distributed training.

Inference

Using a trained AI model to generate predictions or outputs from new input data. Distinct from training (learning model parameters from data). Production AI systems spend most of their compute budget on inference. Inference optimization (quantization, batching, caching) is critical for cost management.

J–M

KV Cache

Key-value cache — stores intermediate attention computation results during LLM inference to avoid recomputing them for each new token. KV cache size limits the maximum context length and concurrent request capacity. PagedAttention (vLLM) manages KV cache memory efficiently to increase throughput.

Knowledge Distillation

Training a smaller "student" model to mimic the behavior of a larger "teacher" model. Produces compact models that retain much of the teacher's performance at a fraction of the serving cost. Widely used to create efficient models for production deployment.

LoRA (Low-Rank Adaptation)

A parameter-efficient fine-tuning technique that adds small trainable matrices to frozen pre-trained model weights, dramatically reducing the number of parameters that need to be updated during fine-tuning. Enables fine-tuning large models on modest GPU hardware. QLoRA combines LoRA with quantization for even lower memory requirements.

LLM (Large Language Model)

A transformer-based neural network trained on large text corpora to understand and generate human language. Examples: GPT-4, Claude, Llama, Mistral. LLMs are the foundation of generative AI applications including chatbots, document intelligence, and code generation.

MLOps

Machine Learning Operations — the practices, tools, and processes for deploying, monitoring, and maintaining ML models in production. Analogous to DevOps for software. Encompasses model training pipelines, experiment tracking, model registry, serving infrastructure, monitoring, and automated retraining.

Model Registry

A centralized repository for storing, versioning, and managing trained ML models. Tracks model metadata (training data, hyperparameters, performance metrics), manages model lifecycle stages (staging, production, archived), and provides access controls. Essential for model governance and reproducibility.

Model Risk Management (MRM)

The systematic process of identifying, measuring, and mitigating risks from AI model errors, misuse, or unexpected behavior. Formalized in US financial services regulation (SR 11-7) and increasingly applied across industries. Core components: model inventory, independent validation, approval process, and ongoing monitoring.

N–P

NVLink

NVIDIA's high-speed GPU-to-GPU interconnect. NVLink 4.0 (used in DGX H100) provides 900 GB/s bidirectional bandwidth between all 8 GPUs in a server. Critical for distributed training performance within a single server. Significantly faster than PCIe for GPU-to-GPU communication.

Perplexity

A metric for evaluating language model quality — measures how well the model predicts a sample of text. Lower perplexity indicates better performance. Used to compare language models on benchmark datasets. Not directly correlated with performance on downstream tasks.

Private AI

AI deployment within the organizational security boundary, with full control over data and models. Data does not leave the organizational perimeter. Can be implemented on-premises, in a dedicated private cloud, or in a colocation facility. Required for regulated industries and organizations with strict data sovereignty requirements.

Prompt Engineering

The practice of crafting input prompts to elicit desired outputs from large language models. Techniques include few-shot prompting (providing examples), chain-of-thought prompting (asking the model to reason step-by-step), and system prompts (setting context and constraints). A lower-cost alternative to fine-tuning for many applications.

Prompt Injection

An attack in which malicious instructions embedded in user inputs or retrieved content override the system prompt, causing the LLM to perform unauthorized actions. A critical security risk for LLM applications with tool-use capabilities. Mitigated through input sanitization, privilege separation, and output filtering.

Q–S

Quantization

Reducing the numerical precision of model weights and activations (e.g., from FP32 to FP16, INT8, or INT4) to reduce memory requirements and increase inference throughput. INT8 quantization typically reduces model size by 2x with minimal accuracy loss. INT4 quantization (GPTQ, AWQ) reduces size by 4x with moderate accuracy impact.

RAG (Retrieval-Augmented Generation)

A technique that combines a language model with a retrieval system — the model retrieves relevant documents from a knowledge base before generating a response. Reduces hallucination, enables up-to-date information, and grounds responses in organizational documents. More cost-effective than fine-tuning for knowledge-intensive applications.

RoCEv2

RDMA over Converged Ethernet version 2 — enables Remote Direct Memory Access over standard Ethernet infrastructure. An alternative to InfiniBand for AI cluster networking. Lower cost than InfiniBand; performance gap has narrowed with 400GbE. Requires careful network configuration (PFC, ECN) for reliable RDMA operation.

SHAP (SHapley Additive exPlanations)

A game-theoretic approach to explaining individual model predictions by computing the contribution of each feature to the prediction. The most widely used post-hoc explainability method. Provides consistent, locally accurate explanations for any model. Computationally expensive for large models.

Speculative Decoding

An LLM inference optimization that uses a smaller draft model to generate candidate tokens that a larger model verifies in parallel. Can improve throughput by 2–3x for latency-sensitive applications without changing model outputs. Requires a compatible draft model and adds implementation complexity.

T–Z

Tensor Parallelism

A distributed training and inference technique that splits individual model layers across multiple GPUs. Required for models too large to fit on a single GPU. Distinct from pipeline parallelism (splitting layers across GPUs sequentially) and data parallelism (replicating the model across GPUs with different data).

TensorRT

NVIDIA's SDK for high-performance deep learning inference. Optimizes trained models for NVIDIA GPU deployment through layer fusion, precision calibration, and kernel auto-tuning. TensorRT-LLM extends these optimizations specifically for large language model inference.

Transformer

A neural network architecture based on self-attention mechanisms, introduced in the 2017 paper "Attention Is All You Need." The foundation of modern large language models (GPT, BERT, Llama) and many computer vision models (ViT). Transformers replaced recurrent neural networks (RNNs) as the dominant architecture for sequence modeling.

Triton Inference Server

NVIDIA's open-source inference serving software. Supports multiple frameworks (TensorRT, PyTorch, TensorFlow, ONNX), dynamic batching, model ensembles, and concurrent model execution. The production standard for enterprise AI inference serving at scale.

vLLM

An open-source LLM inference and serving library optimized for high throughput. Key innovations: PagedAttention (efficient KV cache management), continuous batching, and tensor parallelism. Widely adopted for production LLM serving due to its performance and ease of deployment.

Zero-Trust (AI Security)

A security model applied to AI infrastructure: no implicit trust based on network location, continuous verification of all access, and least-privilege access to all AI resources (training data, model weights, inference APIs, MLOps platforms). The appropriate security architecture for enterprise AI infrastructure given the high value of AI assets.