Strategy Questions

Q: Should we move everything to the cloud?

No. A cloud-first strategy does not mean cloud-only. Evaluate each workload based on cost, performance, compliance, and operational requirements. Some workloads are better on-premises (sustained high-utilization, latency-sensitive, regulated data). The goal is workload optimization, not wholesale cloud migration.

Q: What is the difference between hybrid cloud and multi-cloud?

Hybrid cloud combines on-premises infrastructure with public cloud. Multi-cloud uses services from multiple public cloud providers. Hybrid multi-cloud (the most common enterprise architecture) combines on-premises infrastructure with multiple cloud providers.

Q: How do we avoid cloud vendor lock-in?

Use open-source technologies (Kubernetes, PostgreSQL, Kafka) rather than proprietary cloud services where possible. Containerize applications for portability. Use infrastructure-as-code (Terraform) for consistent deployment across providers. Accept some lock-in for services that provide significant value — the cost of portability often exceeds the cost of lock-in.

Q: What is cloud repatriation and when does it make sense?

Cloud repatriation is moving workloads from public cloud back to on-premises or colocation. It makes sense when: workloads run at sustained high utilization (70%+), on-premises TCO is 30%+ lower over 3 years, latency requirements cannot be met in the cloud, or compliance requirements mandate on-premises processing.

Migration Questions

Q: How long does a cloud migration take?

Cloud migrations consistently take longer than initial estimates. A typical enterprise migration (100–500 applications) takes 18–36 months. Individual application migrations range from days (simple rehost) to months (complex refactoring). Discovery and dependency mapping — often underestimated — typically take 2–3 months for a large environment.

Q: What is the 6Rs framework?

The 6Rs (Rehost, Replatform, Refactor, Repurchase, Retire, Retain) provide a framework for evaluating each application's migration strategy. Rehost (lift-and-shift) is fastest; Refactor delivers the most cloud value but requires the most effort. Most migrations use a mix of strategies.

Q: What causes cloud migrations to fail?

Common failure causes: insufficient discovery (undiscovered dependencies cause failures), skipping the landing zone (security and compliance debt), underestimating data migration complexity, no rollback plan, and inadequate testing before cutover. Most migration failures are preventable with proper planning.

Q: Should we migrate databases to the cloud?

Database migration is the highest-risk migration activity. Managed database services (RDS, Azure SQL, Cloud SQL) reduce operational overhead but require careful migration planning. Consider: data volume (large databases take time to migrate), application compatibility (not all applications work with managed database services), and performance requirements (some databases perform better on-premises).

Cost Questions

Q: Why are our cloud costs higher than expected?

Common causes: oversized instances (migrated workloads are often over-provisioned), idle resources (development environments left running), no reserved capacity (paying on-demand rates for predictable workloads), data egress costs (not modeled in initial estimates), and managed service costs (underestimated in initial TCO analysis).

Q: How much can we save with reserved instances?

Reserved instances and savings plans provide 40–60% savings vs. on-demand for predictable workloads. A 1-year Compute Savings Plan provides 40–66% savings; a 3-year plan provides higher discounts. Organizations that cover 70–80% of baseline compute with reserved capacity typically achieve 30–40% overall compute cost reduction.

Q: What is FinOps?

FinOps (Financial Operations) is a cultural practice that brings financial accountability to cloud spending. It combines engineering, finance, and business to optimize cloud costs while maintaining performance and reliability. Organizations with mature FinOps practices achieve 30–50% lower cloud costs than those without.

Q: Are spot instances reliable enough for production?

Spot instances can be used in production for fault-tolerant workloads: stateless web tiers with auto-scaling, batch processing with checkpointing, and CI/CD pipelines. Not appropriate for stateful applications, databases, or workloads with strict SLA requirements. Use a mix of spot and on-demand instances for production workloads that need both cost efficiency and reliability.

Security Questions

Q: Is the cloud secure?

Cloud security is a shared responsibility. The cloud provider secures the underlying infrastructure (physical security, hypervisor, network). You are responsible for securing your workloads (OS patching, application security, data encryption, access controls). Most cloud security incidents are caused by customer misconfiguration, not provider failures.

Q: How do we manage identity across cloud and on-premises?

Identity federation extends on-premises identity (Active Directory) to cloud environments. Azure AD Connect synchronizes on-premises AD to Azure AD for Microsoft-centric organizations. Okta and Ping Identity provide vendor-neutral federation across multiple clouds and on-premises environments.

Q: What is CSPM?

Cloud Security Posture Management — tools that continuously monitor cloud environments for misconfigurations, compliance violations, and security risks. Leading tools: Wiz, Prisma Cloud, Orca Security. Cloud-native options: AWS Security Hub, Azure Security Center, GCP Security Command Center.

Operations Questions

Q: How do we monitor cloud and on-premises infrastructure together?

Unified observability platforms (Datadog, Dynatrace, New Relic) provide monitoring across cloud and on-premises environments. They ingest metrics, logs, and traces from both environments and provide unified dashboards and alerting. This is essential for hybrid cloud environments where applications span both environments.

Q: What is infrastructure as code and why does it matter?

Infrastructure as code (IaC) manages cloud infrastructure through code (Terraform, Pulumi, CloudFormation) rather than manual configuration. Benefits: consistent, repeatable deployments; version control for infrastructure changes; automated testing; and the ability to recreate environments quickly. IaC is essential for managing cloud infrastructure at scale.

Provider Questions

Q: AWS, Azure, or GCP — which should we choose?

It depends on your workloads and existing investments. AWS for general-purpose workloads and the broadest service catalog. Azure for Microsoft-centric organizations (Active Directory, Office 365, SQL Server). GCP for data analytics and ML workloads. Most large enterprises use multiple providers — choose based on workload requirements, not brand preference.

Q: What is Direct Connect / ExpressRoute?

Dedicated network connections between on-premises infrastructure and cloud providers: AWS Direct Connect, Azure ExpressRoute, Google Cloud Interconnect. Provide lower latency (5–20ms vs. 50–150ms for VPN), consistent bandwidth, and private connectivity that bypasses the public internet. Required for latency-sensitive hybrid workloads.

Q: What cloud compliance certifications should I look for?

Key certifications by industry: Healthcare (HIPAA BAA, HITRUST), Financial services (PCI DSS, SOC 2 Type II), Government (FedRAMP, StateRAMP), General (ISO 27001, SOC 2 Type II). All three major providers (AWS, Azure, GCP) hold most major certifications, but verify for your specific region and service.