AI Threat Model

Enterprise AI systems face threats across three layers: the AI infrastructure layer (servers, networks, storage), the AI platform layer (MLOps tools, model registries, serving infrastructure), and the AI application layer (models themselves and the data they process).

AI-Specific Threats

  • Training data poisoning: Adversaries inject malicious samples into training data to cause the model to learn incorrect behaviors or create backdoors
  • Model extraction: Repeated API queries to reconstruct a proprietary model, stealing intellectual property
  • Adversarial examples: Carefully crafted inputs designed to cause model misclassification
  • Prompt injection: Malicious instructions embedded in user inputs that override system prompts in LLM applications
  • Model inversion: Reconstructing training data from model outputs, potentially exposing sensitive information
  • Membership inference: Determining whether specific data was used in model training, violating privacy

Traditional Threats Applied to AI

AI infrastructure also faces conventional threats: unauthorized access to GPU clusters, data exfiltration from training datasets, supply chain attacks on ML frameworks and pre-trained models, and insider threats from AI engineers with broad data access.

Infrastructure Security

AI infrastructure security follows zero-trust principles: no implicit trust based on network location, continuous verification of all access, and least-privilege access to all resources.

Network Segmentation

AI infrastructure should be segmented into dedicated security zones: training cluster network (isolated from corporate network), inference serving network (controlled external access), storage network (accessible only from authorized compute), and management network (out-of-band, restricted access).

Physical Security

GPU clusters represent significant capital investment and contain sensitive model weights and training data. Physical access controls (biometric authentication, mantrap entry, CCTV, access logging) are required for AI infrastructure areas.

Supply Chain Security

ML frameworks (PyTorch, TensorFlow), pre-trained model weights, and Python packages are common supply chain attack vectors. Controls include: verified package sources, software bill of materials (SBOM), container image scanning, and model provenance verification.

Secure Boot and Firmware

GPU servers should use UEFI Secure Boot, verified firmware, and TPM-based attestation to ensure infrastructure integrity. NVIDIA GPU firmware should be verified against NVIDIA's published signatures.

Training Data Security

Training data is a high-value target — it contains the organizational knowledge that makes AI models valuable. Training data security requires controls across the full data lifecycle.

Access Controls

Training datasets should be accessible only to authorized ML engineers and automated training pipelines. Role-based access control (RBAC) with least-privilege principles, multi-factor authentication for data access, and comprehensive audit logging of all data access.

Data Encryption

Training data should be encrypted at rest (AES-256) and in transit (TLS 1.3). Encryption key management should use a dedicated key management service (KMS) with hardware security module (HSM) backing for the highest-sensitivity datasets.

Data Integrity

Cryptographic hashing of training datasets enables detection of unauthorized modifications (potential poisoning attacks). Hash verification should be performed before each training run.

Sensitive Data Handling

Training data containing PII, PHI, or other sensitive information requires additional controls: differential privacy techniques during training, synthetic data generation to replace sensitive records where possible, and strict retention and deletion policies.

Model Security

Trained model weights represent significant intellectual property and must be protected against theft, unauthorized modification, and extraction.

Model Weight Protection

Model weights stored in model registries should be encrypted at rest, access-controlled with RBAC, and versioned with cryptographic signatures to detect unauthorized modifications.

Model Extraction Defense

Defenses against model extraction attacks include: rate limiting on inference APIs, query monitoring for extraction patterns (high-volume systematic queries), output perturbation (adding noise to outputs without degrading utility), and watermarking model outputs to detect stolen models.

Adversarial Robustness

For high-stakes applications (fraud detection, medical diagnosis, autonomous systems), adversarial robustness testing should be part of the model validation process. Adversarial training (including adversarial examples in training data) improves robustness but increases training cost.

LLM-Specific Security

Large language model applications introduce security risks that do not exist in traditional ML systems.

Prompt Injection

Prompt injection attacks embed malicious instructions in user inputs that override system prompts, causing the LLM to perform unauthorized actions. Defenses include: input sanitization, prompt hardening (clear separation between system and user content), output filtering, and privilege separation (LLMs should not have direct access to sensitive systems).

Indirect Prompt Injection

A more sophisticated variant where malicious instructions are embedded in content the LLM retrieves (web pages, documents, emails) rather than direct user input. Particularly dangerous for LLM agents with tool-use capabilities.

Data Leakage via LLMs

LLMs fine-tuned on sensitive data may leak that data in responses. Differential privacy during fine-tuning and output filtering for sensitive patterns (PII, credentials, proprietary information) are required controls.

RAG Security

Retrieval-augmented generation (RAG) systems retrieve documents from a knowledge base to augment LLM responses. Security requirements: access controls on the knowledge base (users should only retrieve documents they are authorized to see), query logging, and injection detection in retrieved content.

Access Controls

AI infrastructure access control must address multiple layers:

  • Infrastructure access: SSH/RDP to GPU servers, Kubernetes cluster access, storage system access
  • Platform access: MLOps platform, model registry, experiment tracking, feature store
  • Data access: Training datasets, feature stores, inference logs
  • Model access: Model weights, inference APIs, model configuration

All access should be governed by: identity-based authentication (no shared credentials), multi-factor authentication for privileged access, just-in-time access provisioning for sensitive operations, and comprehensive audit logging.

Security Monitoring

AI-specific security monitoring extends traditional SIEM/SOC capabilities:

  • Anomalous inference query patterns (potential model extraction)
  • Unusual training data access patterns
  • Model performance degradation (potential adversarial attack or data poisoning)
  • Prompt injection attempts in LLM applications
  • Unauthorized model weight access or modification
  • GPU cluster access from unauthorized sources

Compliance Alignment

AI security controls must align with applicable compliance frameworks:

  • NIST AI RMF: Risk management framework specifically for AI systems, covering governance, mapping, measurement, and management of AI risks
  • ISO/IEC 42001: AI management system standard with security requirements
  • SOC 2 Type II: Security, availability, and confidentiality controls applicable to AI infrastructure
  • NIST CSF 2.0: Updated cybersecurity framework with AI-specific guidance