Azure Machine Learning Architecture Overview

Azure Machine Learning is a cloud service for accelerating and managing the machine learning project lifecycle. The core infrastructure components are:

Compute
Scalable GPU and CPU compute for training, development, and inference — from single VMs to multi-node clusters.
Storage
Datastores connecting Azure Blob, ADLS, Azure Files, and high-performance parallel file systems for training data.
Networking
VNet integration, private endpoints, and managed VNet for secure, isolated AI workloads.
MLOps
Pipelines, model registry, environments, and deployment infrastructure for production ML lifecycle management.

Compute Options

Azure ML Compute Types: Use Cases and Cost Models

Compute TypeUse CaseGPU OptionsCost ModelBest For
Compute ClustersBatch training, parallel experimentsNC, ND, NV seriesPer-minute when runningTraining, hyperparameter tuning
Compute InstancesInteractive development, notebooksCPU + GPU optionsPer-hour when runningData science development
Serverless ComputeManaged training jobsAuto-selectedPer-jobSimple training without cluster management
Kubernetes (AKS)Production inference servingGPU node poolsPer-node-hourHigh-throughput inference
Managed Online EndpointsReal-time inference APICPU + GPUPer-instance-hourManaged inference with auto-scaling
Batch EndpointsOffline batch inferenceCPU + GPUPer-jobLarge-scale batch scoring

GPU VM Series for AI Workloads

  • NC series (T4, A10): Cost-effective inference and light training — $0.90–$3.00/hr
  • ND series (A100, H100): High-performance training and large model inference — $3.00–$32.00/hr
  • NV series (A10, A40): Visualization and graphics workloads — $1.20–$4.50/hr

Spot VMs for Training Cost Reduction

Azure Spot VMs (equivalent to AWS Spot Instances) offer 60–80% discounts on GPU compute in exchange for the possibility of eviction when Azure needs capacity. For training workloads with checkpoint support, spot VMs dramatically reduce training costs.

AML Compute Clusters support low-priority VMs (Azure's spot equivalent) natively. Configure your training jobs to save checkpoints every 10–30 minutes so they can resume after eviction without losing significant progress.

Storage Architecture

Default Storage: Azure Blob

Azure ML workspaces use Azure Blob Storage as the default datastore. Blob storage is cost-effective and scalable but has throughput limitations for large-scale training:

  • Single blob: up to 60 Gbps read throughput
  • Storage account: up to 200 Gbps aggregate
  • Suitable for: model artifacts, experiment outputs, small-to-medium training datasets

High-Performance Storage for Large-Scale Training

For training workloads that require high-throughput data access (large image datasets, video data, genomics), standard Blob storage creates a data loading bottleneck. Options:

  • Azure NetApp Files: NFS-based, up to 4.5 GiB/s per volume — good for medium-scale training
  • Azure Managed Lustre: Parallel file system, up to 1 TB/s aggregate — for large-scale HPC and AI training
  • Premium Blob (P-series): Higher IOPS and throughput than standard Blob — good middle ground

Storage Tier Strategy

Use a tiered storage strategy: hot Blob or NetApp Files for active training data, cool Blob for recent experiment outputs, archive Blob for long-term dataset retention. This can reduce storage costs by 60–70% compared to keeping all data in hot tier.

AML Datastores

AML datastores abstract the underlying storage, allowing training scripts to reference data by datastore name rather than storage account URL. Register all storage accounts as datastores during workspace setup — this simplifies data access and enables data versioning through AML datasets.

Networking

VNet Integration (Required for Enterprise)

Deploy the AML workspace with VNet integration from day one. Retrofitting VNet integration to an existing workspace requires recreating compute resources. Key components:

  • AML workspace with private endpoint in a dedicated subnet
  • Compute cluster subnet (minimum /24 for large clusters — AML allocates IPs per node)
  • Storage account private endpoints (Blob, File, Queue, Table)
  • Azure Container Registry private endpoint (for custom Docker images)
  • Key Vault private endpoint (for secrets management)

Managed VNet (Simplified Option)

AML Managed VNet automatically creates and manages the network isolation for your workspace — including private endpoints for all associated resources. This simplifies setup but provides less control than a custom VNet. Suitable for organizations without dedicated network engineering resources.

Outbound Traffic

AML compute nodes require outbound access to: Azure Container Registry (for Docker images), Azure Storage (for data and artifacts), Azure Monitor (for logging), and PyPI/conda (for package installation). Configure NSG rules or Azure Firewall to allow these specific destinations while blocking all other outbound traffic.

MLOps Pipeline Architecture

A production MLOps pipeline automates the full ML lifecycle: data preparation, training, evaluation, registration, and deployment. AML provides native components for each stage.

Data
Data Versioning and Preparation
Register datasets in AML Data Assets with versioning. Use AML Pipelines for repeatable data preparation steps. Connect to Azure Data Factory for complex ETL workflows.
Train
Training Pipeline
AML Pipelines orchestrate multi-step training workflows. Each step runs in an isolated environment (Docker container) on specified compute. Support for distributed training across multiple nodes.
Evaluate
Model Evaluation
Automated evaluation against held-out test sets. Compare new model against current production model (champion/challenger). Gate promotion to registry based on metric thresholds.
Register
Model Registry
AML Model Registry stores versioned model artifacts with metadata (training metrics, data lineage, responsible AI metrics). Supports model lifecycle stages: staging, production, archived.
Deploy
Deployment
Managed Online Endpoints for real-time inference with auto-scaling. Batch Endpoints for offline scoring. Blue-green deployment for zero-downtime model updates.
Monitor
Production Monitoring
AML Data Drift monitoring detects when production data diverges from training data. Integration with Azure Monitor for latency, throughput, and error rate alerting.

Cost Optimization

  • Scale compute clusters to zero: Set minimum node count = 0 on all compute clusters. Clusters scale to zero after a configurable idle period (default 120 seconds). This eliminates idle compute costs — the most common source of AML cost waste.
  • Use spot/low-priority VMs: 60–80% discount for fault-tolerant training jobs with checkpoint support.
  • Right-size compute: Profile your training jobs — many are CPU-bound or I/O-bound, not GPU-bound. Use CPU compute for data preparation and preprocessing steps.
  • Reserved Instances: For predictable inference workloads, Azure Reserved VM Instances save up to 72% vs. on-demand pricing over a 3-year term.
  • Storage tiering: Move experiment outputs older than 30 days to cool tier, older than 90 days to archive tier.
  • Monitor with Azure Cost Management: Set budget alerts at 80% and 100% of monthly AML budget. Tag all AML resources with project and team for cost allocation.

Enterprise Deployment Patterns

Hub-and-Spoke AML Architecture

For large enterprises with multiple teams, a hub-and-spoke architecture provides centralized governance with team autonomy:

  • Hub workspace: Shared compute clusters, centralized model registry, shared datastores, governance policies
  • Spoke workspaces: Team-specific workspaces with access to hub resources — isolated experiments, team-specific compute quotas
  • Benefits: Cost sharing for expensive GPU clusters, centralized compliance controls, team-level cost visibility

Dev/Test/Prod Workspace Separation

Maintain separate AML workspaces for development, testing, and production. This provides:

  • Isolation between experimental and production workloads
  • Different security policies per environment (dev can be more permissive)
  • Clear promotion path: model trained in dev → evaluated in test → deployed in prod
  • Cost visibility by environment

Frequently Asked Questions

What compute options does Azure Machine Learning offer?
Azure Machine Learning offers six compute types: Compute Clusters (scalable GPU clusters for training), Compute Instances (single-node VMs for development), Serverless Compute (managed training without cluster management), Kubernetes (AKS) for production inference, Managed Online Endpoints for real-time inference APIs, and Batch Endpoints for large-scale offline scoring.
How do you reduce Azure Machine Learning costs?
Key Azure ML cost optimization strategies: (1) set minimum node count to 0 on compute clusters so they scale to zero when idle, (2) use spot/low-priority VMs for fault-tolerant training jobs (60–80% discount), (3) right-size compute — many training jobs are CPU-bound, not GPU-bound, (4) use storage tiering to move old experiment outputs to cool/archive tier, (5) use Azure Reserved VM Instances for predictable inference workloads.
What storage does Azure Machine Learning use?
Azure Machine Learning uses Azure Blob Storage as its default datastore for training data, model artifacts, and experiment outputs. For high-throughput training workloads, Azure NetApp Files or Azure Managed Lustre provide parallel file system performance. AML datastores abstract the underlying storage, allowing you to register multiple storage accounts and reference them by name in training scripts.