Azure Machine Learning Architecture Overview
Azure Machine Learning is a cloud service for accelerating and managing the machine learning project lifecycle. The core infrastructure components are:
Compute Options
Azure ML Compute Types: Use Cases and Cost Models
| Compute Type | Use Case | GPU Options | Cost Model | Best For |
|---|---|---|---|---|
| Compute Clusters | Batch training, parallel experiments | NC, ND, NV series | Per-minute when running | Training, hyperparameter tuning |
| Compute Instances | Interactive development, notebooks | CPU + GPU options | Per-hour when running | Data science development |
| Serverless Compute | Managed training jobs | Auto-selected | Per-job | Simple training without cluster management |
| Kubernetes (AKS) | Production inference serving | GPU node pools | Per-node-hour | High-throughput inference |
| Managed Online Endpoints | Real-time inference API | CPU + GPU | Per-instance-hour | Managed inference with auto-scaling |
| Batch Endpoints | Offline batch inference | CPU + GPU | Per-job | Large-scale batch scoring |
GPU VM Series for AI Workloads
- NC series (T4, A10): Cost-effective inference and light training — $0.90–$3.00/hr
- ND series (A100, H100): High-performance training and large model inference — $3.00–$32.00/hr
- NV series (A10, A40): Visualization and graphics workloads — $1.20–$4.50/hr
Spot VMs for Training Cost Reduction
Azure Spot VMs (equivalent to AWS Spot Instances) offer 60–80% discounts on GPU compute in exchange for the possibility of eviction when Azure needs capacity. For training workloads with checkpoint support, spot VMs dramatically reduce training costs.
AML Compute Clusters support low-priority VMs (Azure's spot equivalent) natively. Configure your training jobs to save checkpoints every 10–30 minutes so they can resume after eviction without losing significant progress.
Storage Architecture
Default Storage: Azure Blob
Azure ML workspaces use Azure Blob Storage as the default datastore. Blob storage is cost-effective and scalable but has throughput limitations for large-scale training:
- Single blob: up to 60 Gbps read throughput
- Storage account: up to 200 Gbps aggregate
- Suitable for: model artifacts, experiment outputs, small-to-medium training datasets
High-Performance Storage for Large-Scale Training
For training workloads that require high-throughput data access (large image datasets, video data, genomics), standard Blob storage creates a data loading bottleneck. Options:
- Azure NetApp Files: NFS-based, up to 4.5 GiB/s per volume — good for medium-scale training
- Azure Managed Lustre: Parallel file system, up to 1 TB/s aggregate — for large-scale HPC and AI training
- Premium Blob (P-series): Higher IOPS and throughput than standard Blob — good middle ground
Storage Tier Strategy
AML Datastores
AML datastores abstract the underlying storage, allowing training scripts to reference data by datastore name rather than storage account URL. Register all storage accounts as datastores during workspace setup — this simplifies data access and enables data versioning through AML datasets.
Networking
VNet Integration (Required for Enterprise)
Deploy the AML workspace with VNet integration from day one. Retrofitting VNet integration to an existing workspace requires recreating compute resources. Key components:
- AML workspace with private endpoint in a dedicated subnet
- Compute cluster subnet (minimum /24 for large clusters — AML allocates IPs per node)
- Storage account private endpoints (Blob, File, Queue, Table)
- Azure Container Registry private endpoint (for custom Docker images)
- Key Vault private endpoint (for secrets management)
Managed VNet (Simplified Option)
AML Managed VNet automatically creates and manages the network isolation for your workspace — including private endpoints for all associated resources. This simplifies setup but provides less control than a custom VNet. Suitable for organizations without dedicated network engineering resources.
Outbound Traffic
AML compute nodes require outbound access to: Azure Container Registry (for Docker images), Azure Storage (for data and artifacts), Azure Monitor (for logging), and PyPI/conda (for package installation). Configure NSG rules or Azure Firewall to allow these specific destinations while blocking all other outbound traffic.
MLOps Pipeline Architecture
A production MLOps pipeline automates the full ML lifecycle: data preparation, training, evaluation, registration, and deployment. AML provides native components for each stage.
Cost Optimization
- Scale compute clusters to zero: Set minimum node count = 0 on all compute clusters. Clusters scale to zero after a configurable idle period (default 120 seconds). This eliminates idle compute costs — the most common source of AML cost waste.
- Use spot/low-priority VMs: 60–80% discount for fault-tolerant training jobs with checkpoint support.
- Right-size compute: Profile your training jobs — many are CPU-bound or I/O-bound, not GPU-bound. Use CPU compute for data preparation and preprocessing steps.
- Reserved Instances: For predictable inference workloads, Azure Reserved VM Instances save up to 72% vs. on-demand pricing over a 3-year term.
- Storage tiering: Move experiment outputs older than 30 days to cool tier, older than 90 days to archive tier.
- Monitor with Azure Cost Management: Set budget alerts at 80% and 100% of monthly AML budget. Tag all AML resources with project and team for cost allocation.
Enterprise Deployment Patterns
Hub-and-Spoke AML Architecture
For large enterprises with multiple teams, a hub-and-spoke architecture provides centralized governance with team autonomy:
- Hub workspace: Shared compute clusters, centralized model registry, shared datastores, governance policies
- Spoke workspaces: Team-specific workspaces with access to hub resources — isolated experiments, team-specific compute quotas
- Benefits: Cost sharing for expensive GPU clusters, centralized compliance controls, team-level cost visibility
Dev/Test/Prod Workspace Separation
Maintain separate AML workspaces for development, testing, and production. This provides:
- Isolation between experimental and production workloads
- Different security policies per environment (dev can be more permissive)
- Clear promotion path: model trained in dev → evaluated in test → deployed in prod
- Cost visibility by environment