Waste Elimination
Waste elimination is the fastest path to cloud cost reduction. Common sources of cloud waste:
Idle Resources
Instances running at less than 5% CPU utilization for extended periods. Common causes: development environments left running, forgotten test instances, and over-provisioned infrastructure. Identify with cloud cost management tools; terminate or stop idle instances. Typical savings: 10–20% of compute spend.
Unattached Storage
EBS volumes, Azure Managed Disks, and GCP Persistent Disks that are not attached to any instance. Created when instances are terminated without deleting their storage. Identify with cloud cost management tools; delete unneeded volumes after verifying no data is needed. Typical savings: 5–10% of storage spend.
Unused Reserved Capacity
Reserved instances or savings plans that are not being utilized. Occurs when workloads are terminated or migrated but reservations are not cancelled. Sell unused AWS Reserved Instances on the Reserved Instance Marketplace; convert to more flexible savings plans.
Orphaned Snapshots
Storage snapshots for instances that no longer exist. Accumulate over time and can represent significant storage costs. Implement snapshot lifecycle policies to automatically delete old snapshots.
Oversized Load Balancers and NAT Gateways
Load balancers and NAT gateways that are provisioned but carry minimal traffic. Review and consolidate where possible.
Rightsizing
Rightsizing matches instance size to actual workload requirements. Most workloads migrated to the cloud are oversized — organizations provision conservatively to avoid performance issues, then never revisit sizing.
Rightsizing Process
- Collect 2–4 weeks of CPU, memory, network, and storage utilization data
- Identify instances where peak utilization is consistently below 40% of provisioned capacity
- Recommend smaller instance types that provide sufficient headroom (target 60–70% peak utilization)
- Test in non-production environment before applying to production
- Apply changes during maintenance windows with rollback plan
- Monitor for 2 weeks post-change to verify performance
Rightsizing Tools
AWS Compute Optimizer, Azure Advisor, and GCP Recommender provide automated rightsizing recommendations. Third-party tools (CloudHealth, Spot.io) provide more sophisticated recommendations across multiple providers.
Typical Savings
Rightsizing typically achieves 20–30% reduction in compute costs. Combined with waste elimination, organizations commonly achieve 30–40% total cost reduction in the first 90 days of a FinOps program.
Reserved Capacity
Reserved capacity provides significant discounts for committing to use specific resources for 1 or 3 years. The most impactful optimization for predictable workloads.
AWS Reserved Instances and Savings Plans
- Standard Reserved Instances: 40–60% discount vs. on-demand for specific instance type in specific region. Least flexible; highest discount.
- Convertible Reserved Instances: 30–45% discount; can exchange for different instance type. More flexible than standard.
- Compute Savings Plans: 40–66% discount; applies to any EC2 instance, Fargate, and Lambda. Most flexible; applies automatically.
- EC2 Instance Savings Plans: 60–72% discount; applies to specific instance family in specific region. Less flexible than Compute Savings Plans but higher discount.
Azure Reserved VM Instances
40–72% discount vs. pay-as-you-go for 1 or 3 year commitment. Can be exchanged or cancelled (with fee). Applies to specific VM size in specific region.
Commitment Strategy
Cover 70–80% of baseline compute with reserved capacity; leave 20–30% on-demand for flexibility. Use 1-year commitments initially to maintain flexibility; move to 3-year for stable, long-term workloads. Review and adjust commitments quarterly.
Spot & Preemptible Instances
Spot instances (AWS), Spot VMs (Azure), and Preemptible VMs (GCP) provide access to unused cloud capacity at 60–80% discount vs. on-demand. The tradeoff: instances can be terminated with 2-minute notice when the provider needs the capacity back.
Appropriate Use Cases
- Batch processing jobs that can be checkpointed and restarted
- AI/ML training runs with checkpoint support
- CI/CD pipelines
- Stateless web tier with auto-scaling
- Data processing and analytics
Inappropriate Use Cases
- Stateful applications that cannot tolerate interruption
- Databases (unless using managed services with automatic failover)
- Applications with strict SLA requirements
Spot Instance Best Practices
- Use multiple instance types and availability zones to reduce interruption probability
- Implement checkpointing for long-running jobs
- Use Spot Instance interruption notices to gracefully handle termination
- Combine with on-demand instances for minimum capacity guarantee
Storage Optimization
Storage costs are often overlooked but can represent 20–30% of total cloud spend. Key optimization opportunities:
Storage Tiering
Move infrequently accessed data to lower-cost storage tiers: AWS S3 Intelligent-Tiering (automatic), S3 Glacier (archival), Azure Cool/Archive tiers, GCP Nearline/Coldline/Archive. Typical savings: 50–80% for data that can tolerate higher retrieval latency.
Snapshot Management
Implement snapshot lifecycle policies to automatically delete old snapshots. Retain only the snapshots needed for recovery (typically 7–30 days for daily snapshots). Snapshots accumulate quickly and can represent significant costs without lifecycle management.
Data Transfer Optimization
Minimize data egress costs by: keeping data and compute in the same region, using CloudFront/CDN for content delivery, compressing data before transfer, and using Direct Connect/ExpressRoute for large data transfers (often cheaper than internet egress).
Database Storage
Review database storage allocation — many databases are provisioned with significantly more storage than needed. Enable storage auto-scaling to avoid over-provisioning. Archive or delete old data that is no longer needed.
Architecture Optimization
Serverless for Variable Workloads
AWS Lambda, Azure Functions, and GCP Cloud Functions charge only for actual execution time. For workloads with variable or infrequent execution, serverless can be 90%+ cheaper than always-on instances.
Managed Services
Replacing self-managed services with managed equivalents (RDS vs. self-managed MySQL, ElastiCache vs. self-managed Redis) reduces operational overhead and often reduces cost when operational labor is included in the comparison.
Auto-Scaling
Implement auto-scaling for all variable workloads. Scale down during off-peak hours (nights, weekends) for workloads with predictable patterns. Typical savings: 20–40% for workloads with significant off-peak periods.
FinOps Practices
FinOps is a cultural practice that brings financial accountability to cloud spending. Key practices:
- Tagging and cost allocation: Tag all resources; allocate costs to business units and applications
- Budgets and alerts: Set budgets for each team/application; alert when spending exceeds thresholds
- Regular cost reviews: Weekly or monthly reviews of cloud costs by team and application
- Chargeback/showback: Make teams accountable for their cloud costs
- Optimization targets: Set specific cost reduction targets for each team
- FinOps team: Dedicated team or role responsible for cloud cost optimization
Organizations with mature FinOps practices consistently achieve 30–50% lower cloud costs than organizations without them. The cultural change — making engineers accountable for the cost of the infrastructure they provision — is more impactful than any specific technical optimization.