Running on two public clouds is easy to start and hard to make useful. Deploying the same container images to two Kubernetes services takes an afternoon. Making the second cloud able to take over when the first one fails, with current data, working credentials, reachable dependencies and a tested runbook, takes months, and most of that work is in layers that containers do not touch.

This article is about building it, not deciding whether to; the strategic case is covered in multi-cloud strategy. It takes the most common serious design, an active-passive service with a primary in one cloud and a warm standby in another, and works through each layer in the order that dependencies force: topology, identity, network, data, traffic steering, the platform layer and the failover runbook. A worked example puts numbers on recovery point, recovery time and the cross-cloud data transfer bill.

Active-passive across two clouds: what has to exist on both sidesTraffic steeringhealth-checked DNS or global load balancerCloud A (primary region)Cloud B (standby region)100%0% until failoverKubernetes clustersame images, same manifestsKubernetes clusterscaled down, warmPostgres primarywritesObject storesource bucketPostgres replicaasync, lag = RPOObject storereplicated copyWAL stream over private linkLanding zone Anative IAM, KMS, policy, loggingLanding zone Bnative IAM, KMS, policy, loggingIdentity providerSSO for people, OIDC for workloadsPrivate interconnectnon-overlapping CIDRs, BGPTelemetry pipelineOpenTelemetry to one backendKubernetes makes compute portable; data, identity and network decide whether failover works.
A warm standby in a second cloud needs replicated data, federated identity, a private network path and traffic steering, not just a second cluster.

Choose the topology from recovery objectives

Name the topology first, because each one implies different data machinery. In a partitioned design each workload or customer segment lives in exactly one cloud; there is no cross-cloud failover and little cross-cloud traffic, which is often what an acquisition or a data-residency rule actually needs. In active-passive one cloud serves all traffic and the other holds replicated data and a scaled-down copy of the service. In active-active both clouds take writes, which requires a database that replicates across them with conflict handling or consensus, and pays the inter-cloud round trip on every consistent write.

TopologyRecovery pointRecovery timeHard part
PartitionedNone (no failover)Same as single cloudShared identity, billing and tooling
Active-passiveReplication lag at failureMinutes, mostly detection and promotionKeeping the standby honest
Active-activeNear zero with synchronous consensusSecondsWrite latency, conflicts, cost

Write the target as two numbers per service, recovery point objective (how much acknowledged data you can lose) and recovery time objective (how long you can be down), and let them choose the topology. If the targets are satisfied by two regions of one provider, that is cheaper and simpler; read multi-region architecture first. Cross-cloud standby earns its cost when the risk you are covering is the provider itself: a control-plane outage spanning regions, an account suspension, or a regulatory requirement for provider exit.

Choose the topology from recovery objectives

Name the topology first, because each one implies different data machinery. In a partitioned design each workload or customer segment lives in exactly one cloud; there is no cross-cloud failover and little cross-cloud traffic, which is often what an acquisition or a data-residency rule actually needs. In active-passive one cloud serves all traffic and the other holds replicated data and a scaled-down copy of the service. In active-active both clouds take writes, which requires a database that replicates across them with conflict handling or consensus, and pays the inter-cloud round trip on every consistent write.

TopologyRecovery pointRecovery timeHard part
PartitionedNone (no failover)Same as single cloudShared identity, billing and tooling
Active-passiveReplication lag at failureMinutes, mostly detection and promotionKeeping the standby honest
Active-activeNear zero with synchronous consensusSecondsWrite latency, conflicts, cost

Write the target as two numbers per service, recovery point objective (how much acknowledged data you can lose) and recovery time objective (how long you can be down), and let them choose the topology. If the targets are satisfied by two regions of one provider, that is cheaper and simpler; read multi-region architecture first. Cross-cloud standby earns its cost when the risk you are covering is the provider itself: a control-plane outage spanning regions, an account suspension, or a regulatory requirement for provider exit.

Identity: one root of trust, native permissions

Each cloud keeps its own IAM, and that is fine: a common abstraction over three permission models usually ends up weaker than any of them. What must be shared is the root of trust. People sign in through one identity provider over SAML or OIDC and receive short-lived roles in each cloud, so disabling a leaver in one place removes access everywhere.

Workloads should hold no long-lived cloud keys at all. All three large providers accept OIDC tokens from an external issuer and exchange them for short-lived credentials: AWS with AssumeRoleWithWebIdentity, Google Cloud with Workload Identity Federation, and Azure with federated identity credentials on a managed identity or app registration. A service running in cloud A presents its platform-issued token to cloud B and gets scoped credentials there. The same mechanism serves CI pipelines. For example, letting deployments from one repository's main branch assume a role in AWS and impersonate a service account in Google Cloud. First the AWS IAM role trust policy, which only main-branch runs of acme/payments can satisfy:

{
  "Version": "2012-10-17",
  "Statement": [{
    "Effect": "Allow",
    "Principal": {"Federated": "arn:aws:iam::111122223333:oidc-provider/token.actions.githubusercontent.com"},
    "Action": "sts:AssumeRoleWithWebIdentity",
    "Condition": {
      "StringEquals": {
        "token.actions.githubusercontent.com:aud": "sts.amazonaws.com",
        "token.actions.githubusercontent.com:sub": "repo:acme/payments:ref:refs/heads/main"
      }
    }
  }]
}
# Google Cloud: a workload identity pool and provider trusting the same issuer
gcloud iam workload-identity-pools create ci-pool --location=global
gcloud iam workload-identity-pools providers create-oidc github \
  --location=global --workload-identity-pool=ci-pool \
  --issuer-uri=https://token.actions.githubusercontent.com \
  --attribute-mapping=google.subject=assertion.sub,attribute.repository=assertion.repository \
  --attribute-condition="assertion.repository=='acme/payments'"

Encryption keys stay native: each cloud's data is encrypted with that cloud's key service, and the standby never depends on the primary's key service to decrypt anything. A failover plan that needs cloud A's KMS to read cloud B's replica has a hidden single point of failure.

Network: address plan, private paths and DNS

Start with an address plan. Every VPC and VNet in every cloud, plus on-premises ranges, needs a non-overlapping CIDR block reserved centrally, because overlapping ranges make routing between them impossible without NAT, and NAT breaks the database replication and service discovery you are about to build. Allocate generously; resizing a production network later is a migration project.

Then choose the path between clouds. IPsec VPN over the internet, with BGP for route exchange, is quick to set up and adequate for control traffic and modest replication. For sustained replication, private connectivity is the norm: colocation fabrics such as Equinix or Megaport that cross-connect to each provider's dedicated interconnect, or provider-built links. Google Cloud offers Cross-Cloud Interconnect to other clouds, Oracle and Microsoft run a direct interconnect between OCI and Azure, and in late 2025 AWS announced AWS Interconnect - multicloud in preview, with Google Cloud as the first partner and Azure planned. Check current availability for your regions before designing around any of them.

Pair regions that are physically close, so the interconnect round trip is a few milliseconds and replication lag stays small. Run the link in redundant pairs on separate devices, and put DNS on the plan too: each cloud's private zones must resolve the other side's service names, through conditional forwarding or a shared resolver, or the standby's services will not find their dependencies after failover. Cloud networking fundamentals covers the VPC side in more detail.

Data: replication lag is your recovery point

Data decides your recovery point. With asynchronous replication, the data you lose in a failover is whatever the primary acknowledged but the replica had not yet applied, so RPO equals replication lag at the moment of failure. That makes lag the most important metric in the whole design: alert on it, chart it next to write throughput, and know its p99 under peak load, not its average on a quiet day.

The mechanism depends on the store. PostgreSQL logical or physical streaming replication to a self-managed replica or to a managed service that accepts an external source works across clouds; check that your managed service supports the direction you need, since some accept inbound replication only during migration. Object storage replication between providers is not built in, so you run a copy job or event-driven sync, with its own lag and its own monitoring. Kafka-style logs can be mirrored with MirrorMaker 2 or equivalent. Each stream needs an owner and a lag alert.

Every byte replicated is cross-cloud egress, charged by the source provider at internet or interconnect rates that are far higher than in-region traffic. Model it per stream with your contracted prices:

# Monthly cross-cloud replication bill, per stream (use your negotiated rates)
def monthly_egress_cost(change_gb_per_day, overhead=1.3, price_per_gb=0.0, port_fee_month=0.0):
    gb = change_gb_per_day * 30 * overhead       # headroom for protocol overhead, retries and growth
    return gb * price_per_gb + port_fee_month

# A 40 GB/day change stream at an assumed $0.02/GB over a private link:
#   40 * 30 * 1.3 = 1,560 GB  ->  $31 transfer + port/cross-connect fees
# The same stream at an assumed $0.09/GB over the internet: ~$140 transfer.
# Full re-seeds after a broken replica multiply this by the database size.

The numbers above are placeholders for the arithmetic, not current list prices. The pattern holds regardless: replication of changes is usually affordable, while designs that move whole datasets repeatedly, or let services in one cloud read chatty APIs in the other, are not. Egress cost in depth explains the pricing models.

Traffic steering and the failover runbook

Clients need a way to arrive at whichever cloud is live. The common choice is health-checked DNS at a provider-neutral DNS service or a global load balancer that can route to origins in both clouds. Keep record TTLs short, 30 to 60 seconds, and still plan for stragglers: some resolvers and client libraries cache longer than the TTL, so the old side must fail closed rather than accept writes after the switch.

That last point is the heart of the runbook. Automatic failover of a stateful primary across clouds is risky because a network partition looks exactly like an outage, and two writable primaries is worse than a short outage. Most teams automate detection and every step of the switch, then keep a human decision before it:

def fail_over_to_b(incident):
    assert incident.declared_by_human
    fence_primary_a()                    # revoke app write role / set DB read-only / scale writers to 0
    lag = replica_b.replay_lag_seconds() # record what you are about to lose
    replica_b.wait_until_caught_up(timeout_s=60, allow_partial=True)
    replica_b.promote()                  # becomes a writable primary
    platform_b.scale(deployments="all", to="production_replicas")
    secrets_b.point_apps_at(replica_b)
    dns.set_weights({"cloud_a": 0, "cloud_b": 100})
    smoke_test(region="b")               # synthetic login, read, write, payment authorisation
    incident.record(rpo_seconds=lag, rto_seconds=incident.elapsed())
    # failback is a separate, scheduled job: rebuild A as a replica of B, then switch in a quiet window

The platform layer: pipelines, landing zones, telemetry

Above the network and data, the platform layer keeps the two sides from drifting. Infrastructure as code should use per-cloud modules behind a shared interface (a service gets a database, a bucket, a queue and credentials) rather than a lowest-common-denominator abstraction that hides each provider's strengths. Kubernetes manifests, built images and policy-as-code rules are shared artifacts promoted to both sides by the same pipeline, so the standby runs exactly the release the primary runs.

Each cloud gets its own landing zone with equivalent guardrails, mapped to one control framework so auditors see one set of controls with two implementations. Telemetry should be emitted through OpenTelemetry and shipped to a single backend outside both failure domains, or duplicated to both, so you can still see the system when one cloud is the problem. Cost tags must use the same keys on both sides so spend rolls up by service, not by provider.

Worked example: a payments API across AWS and Google Cloud

A payments API runs on Kubernetes in AWS Frankfurt with a PostgreSQL primary that generates roughly 40 GB of WAL a day, with short peaks around 4 MB/s. The standby is in Google Cloud Frankfurt, connected by a redundant pair of private links with a 2 to 3 ms round trip. Targets: RPO under 30 seconds, RTO under 15 minutes.

Replication lag measured over a month sits at 300 ms at p50 and 4 seconds at p99, with spikes to 20 seconds during nightly batch jobs, so the 30-second RPO holds with margin. The first quarterly drill took 41 minutes, mostly because the standby cluster had two nodes and needed 25 minutes to scale, and because a secret still pointed at the old database host. The team kept the standby at a third of production capacity with the autoscaler pre-warmed, moved every endpoint into secrets managed by the pipeline, and scripted the smoke test. The second drill completed in 11 minutes with 2 seconds of measured data loss, and the third found that a fraud-scoring dependency still called an API only reachable from AWS, which was the most valuable finding of the year.

Failure modes

  • Split brain. The old primary keeps accepting writes from clients with stale DNS. Fence before you promote, every time.
  • Stale standby. Images, configuration or schema drift because the standby is not deployed by the same pipeline. Deploy both sides on every release.
  • Hidden single-cloud dependencies. A SaaS allow-list, a KMS key, a container registry or a third-party API reachable only from the primary. Only drills find these.
  • Replication silently broken. A stream stops and nobody notices until failover. Alert on lag and on the absence of lag data.
  • Surprise egress. A service in one cloud reads a database in the other. Watch cross-cloud bytes per service in cost reports.

Trade-offs

The trade is money and engineering attention for protection against a narrow but severe class of failure. A warm standby roughly doubles the data tier, adds a fraction of compute, interconnect fees and egress, and needs a team that drills quarterly. Partitioned multi-cloud costs far less but protects nothing at runtime. Active-active buys seconds of RTO at the price of a database designed for it and slower consistent writes. Choose the cheapest topology that meets the written objectives, and revisit it when they change.

What to do next

  1. Write RPO and RTO per service and confirm a second cloud is actually needed to meet them.
  2. Reserve a non-overlapping address plan across every cloud and on-premises network before building anything.
  3. Federate human access through one identity provider and replace long-lived workload keys with OIDC federation.
  4. Stand up the private link and cross-cloud DNS, then start replication and alert on lag and on missing lag data.
  5. Deploy both sides from the same pipeline on every release, with equivalent landing-zone guardrails.
  6. Script the fence, promote, scale, repoint and switch steps, then run a drill and record measured RPO and RTO.
  7. Review cross-cloud egress per service each month alongside your FinOps reporting.
Key takeaway: Multi-cloud resilience is built from the bottom up: pick a topology from written recovery objectives, share one identity root while keeping native permissions and keys, plan addresses and private links before workloads, treat replication lag as your recovery point, fence before you promote, deploy both sides from one pipeline, and prove the whole thing with timed drills.