Network Architecture

Connectivity Options

The network connection between on-premises and cloud environments is the foundation of hybrid cloud architecture. Three options:

  • VPN (IPsec): Encrypted tunnel over the public internet. Lowest cost; highest latency (50–150ms typical); variable bandwidth. Appropriate for non-latency-sensitive workloads and DR scenarios.
  • Dedicated connectivity: AWS Direct Connect, Azure ExpressRoute, Google Cloud Interconnect. Private connection bypassing the public internet. Lower latency (5–20ms), consistent bandwidth, higher cost. Required for latency-sensitive workloads.
  • SD-WAN: Software-defined WAN that can use multiple connections (MPLS, broadband, LTE) with intelligent traffic routing. Provides flexibility and redundancy at lower cost than dedicated connectivity.

Network Topology

Hub-and-spoke topology: on-premises network as the hub, cloud VPCs/VNets as spokes connected via transit gateway or virtual WAN. Provides centralized routing and security inspection. Alternatively, mesh topology for organizations with multiple cloud regions and on-premises sites.

DNS Architecture

Consistent DNS resolution across environments is critical. Options: split-horizon DNS (different responses for on-premises vs. cloud queries), centralized DNS with conditional forwarding, or cloud-native DNS with on-premises forwarding. Azure Private DNS Zones and AWS Route 53 Resolver support hybrid DNS architectures.

Bandwidth Planning

Calculate bandwidth requirements based on: data replication volumes, backup traffic, application traffic between tiers, and monitoring/management traffic. Add 50% headroom for growth. Bandwidth constraints are a common cause of hybrid cloud performance problems.

Identity & Access Management

Consistent identity across on-premises and cloud environments is the foundation of hybrid cloud security. Users should authenticate once and access resources in both environments without separate credentials.

Identity Federation

Federation extends on-premises identity (Active Directory) to cloud environments. Common approaches:

  • Azure AD Connect / Entra ID: Synchronizes on-premises Active Directory to Azure AD. Enables single sign-on to Azure, Microsoft 365, and thousands of SaaS applications. The most common identity federation approach for Microsoft-centric organizations.
  • AWS IAM Identity Center: Connects to Active Directory or third-party identity providers. Enables single sign-on to AWS accounts and applications.
  • Third-party IdP: Okta, Ping Identity, or ForgeRock as a central identity provider for all environments. Provides vendor-neutral identity federation.

Privileged Access Management

Privileged access to cloud infrastructure should be managed through a PAM (Privileged Access Management) solution that provides: just-in-time access, session recording, and approval workflows. CyberArk, BeyondTrust, and cloud-native solutions (AWS IAM, Azure PIM) are common choices.

Service Accounts and Workload Identity

Applications running in hybrid environments need identities to access resources in both environments. Use workload identity federation (AWS IAM Roles Anywhere, Azure Managed Identity) to avoid storing long-lived credentials in application configurations.

Data Architecture

Data architecture is often the most complex aspect of hybrid cloud design. Data gravity — the tendency for applications to stay close to their data — is the most underestimated constraint.

Data Classification and Placement

Classify data by sensitivity and compliance requirements to determine placement:

  • Regulated data (PHI, PII, financial): May require on-premises storage; verify cloud compliance certifications before moving
  • Sensitive business data: On-premises or private cloud; encrypted in transit and at rest
  • Operational data: Can reside in cloud; optimize for cost and performance
  • Public data: Cloud-native storage; optimize for global distribution and cost

Data Synchronization

Applications that span environments need data synchronization. Patterns:

  • Database replication: Replicate on-premises databases to cloud (AWS DMS, Azure Database Migration Service)
  • Event streaming: Apache Kafka or cloud-native event streaming for real-time data synchronization
  • ETL/ELT pipelines: Batch data movement for analytics workloads
  • Object storage sync: AWS S3 Transfer Acceleration, Azure Data Box for large dataset transfers

Data Egress Costs

Cloud providers charge for data egress (data leaving the cloud). For data-intensive hybrid workloads, egress costs can be significant. Design data flows to minimize egress: process data where it lives, use cloud-native analytics for cloud data, and avoid unnecessary data movement between environments.

Workload Placement

Workload placement decisions should be driven by data location, compliance requirements, performance requirements, and cost:

On-Premises Workloads

  • Applications with strict data sovereignty requirements
  • Latency-sensitive applications that require sub-5ms access to on-premises data
  • Applications with sustained high compute requirements (lower TCO on-premises)
  • Legacy applications that cannot be easily modified for cloud deployment

Cloud Workloads

  • Variable workloads that benefit from elastic scaling
  • Development and test environments
  • Disaster recovery and backup
  • Global applications requiring geographic distribution
  • Applications that benefit from cloud-native services (AI/ML, analytics, IoT)

Workload Portability

Design workloads for portability where possible: containerization (Docker, Kubernetes) enables workloads to run in any environment; infrastructure-as-code (Terraform, Pulumi) enables consistent deployment across environments; cloud-agnostic APIs reduce vendor lock-in.

Security Architecture

Zero-Trust for Hybrid Cloud

Zero-trust architecture is the appropriate security model for hybrid cloud: no implicit trust based on network location, continuous verification of all access, and least-privilege access to all resources. Key components: identity verification (MFA, conditional access), device compliance, network micro-segmentation, and application-level access controls.

Security Information and Event Management (SIEM)

Centralized security monitoring across on-premises and cloud environments. Cloud-native SIEM (Microsoft Sentinel, AWS Security Hub) can ingest logs from both environments. On-premises SIEM (Splunk, IBM QRadar) can be extended to cloud environments.

Data Protection

Consistent data protection across environments: encryption at rest (AES-256) and in transit (TLS 1.3), key management (on-premises HSM or cloud KMS), and data loss prevention (DLP) policies that apply across environments.

Management & Operations

Infrastructure as Code

Manage hybrid cloud infrastructure through code: Terraform for multi-cloud infrastructure provisioning, Ansible for configuration management, and GitOps workflows for change management. Infrastructure-as-code enables consistent, repeatable deployments across environments.

Observability

Unified observability across environments: metrics (Prometheus, Datadog), logs (Elasticsearch, Splunk), and traces (Jaeger, AWS X-Ray). Distributed tracing is essential for debugging performance issues in applications that span environments.

Cost Management

FinOps practices for hybrid cloud: tag all cloud resources for cost attribution, set budgets and alerts, regularly review and rightsize cloud resources, and compare cloud vs. on-premises costs for workloads to identify repatriation candidates.

Reference Architectures

Regulated Industry Hybrid Cloud

Sensitive data and regulated workloads on-premises; non-regulated workloads in the cloud. Dedicated connectivity (Direct Connect/ExpressRoute). Identity federation via Azure AD Connect or Okta. Centralized security monitoring via SIEM. Data classification enforced by DLP policies.

Cloud-Bursting Architecture

Primary workloads on-premises; burst capacity in the cloud. Kubernetes federation (on-premises cluster + cloud cluster) for workload portability. Shared storage via cloud-native NFS or object storage. Auto-scaling policies that trigger cloud bursting when on-premises capacity is exhausted.

Hybrid DR Architecture

Primary workloads on-premises; DR in the cloud. Continuous replication via AWS CloudEndure or Azure Site Recovery. RTO of 1–4 hours; RPO of minutes. Regular DR testing (quarterly) to verify recovery procedures.