Running on Google Cloud in more than one region protects you from the failure of an entire region and puts your service closer to users. Google gives you unusually strong building blocks for it: a single anycast load balancer IP that fronts backends in many regions, and Spanner, a database that commits synchronously across regions. It also gives you services whose "multi-region" label means something much weaker than you might assume.

This article walks through a reference architecture, then the property that matters most for each data service, Spanner placement and read choices in code, a worked example with numbers, the failover sequence, and the failure modes. Product behaviour was checked against Google Cloud documentation in October 2026; re-check anything you build on, because these services change.

A multi-region service on Google CloudUsersone anycast IPGlobal external Application LBedge TLS, Cloud Armor, URL mapRegion us-east4Cloud Run / GKEstatelessMemorystoreregional cacheSpanner read-write replicas [L]Region us-east1Cloud Run / GKEstatelessMemorystoreregional cacheSpanner read-write replicasnearest healthyoverflow / failoverWitness us-central1votes, no full datasync PaxosGCS dual-region bucketasync, turbo RPO 15 minPub/Sub (global service)storage policy per topicCompute is duplicated per region; the data tier decides RPO, RTO and write latency.
One anycast load balancer in front of regional compute; Spanner commits synchronously across two read-write regions plus a witness, while storage and caches replicate asynchronously or not at all.

What multi-region means on Google Cloud

On Google Cloud, location is a property of each resource. A zone is one failure domain inside a region; a regional resource survives a zone failure; a multi-region or dual-region resource is meant to survive a region failure. Your application is only as multi-region as its least multi-region dependency, so the first step is to list every dependency and its real location type.

Stateless compute is the easy part: deploy the same container to Cloud Run or GKE in two or three regions. The hard part is state, because every replication scheme chooses between synchronous replication, which costs write latency, and asynchronous replication, which costs data on failover (the recovery point objective, RPO). The general patterns are covered in cloud multi-region architecture; here we stay with what Google's services actually do.

What multi-region means on Google Cloud

On Google Cloud, location is a property of each resource. A zone is one failure domain inside a region; a regional resource survives a zone failure; a multi-region or dual-region resource is meant to survive a region failure. Your application is only as multi-region as its least multi-region dependency, so the first step is to list every dependency and its real location type.

Stateless compute is the easy part: deploy the same container to Cloud Run or GKE in two or three regions. The hard part is state, because every replication scheme chooses between synchronous replication, which costs write latency, and asynchronous replication, which costs data on failover (the recovery point objective, RPO). The general patterns are covered in cloud multi-region architecture; here we stay with what Google's services actually do.

Choosing a topology

Before picking services, pick a topology for each workload. There are three common shapes, and one system often uses more than one.

TopologyTrafficDataTypical RPO / RTO
Active-activeAll regions serve reads and writesSynchronous (Spanner, Firestore)Zero / seconds, automatic
Active-active reads, single writerAll regions read, one region writesAsync replicas for reads, primary in one regionSeconds to minutes / minutes, promotion needed
Warm standbyOne region serves, the other is scaled downAsync replication, backupsMinutes / tens of minutes, runbook driven

Active-active is the only shape in which a region loss needs no human decision, and on Google Cloud it is practical mainly because Spanner and Firestore make the data tier synchronous. The single-writer shape fits a Cloud SQL or AlloyDB primary with cross-region replicas: reads stay local, but every write crosses to the primary region, and losing that region means promoting a replica. Warm standby is the cheapest and the slowest, and its weak point is that a scaled-down region is rarely tested at full load.

Whatever you choose, an asynchronous promotion follows the same order, and getting the order wrong causes split brain:

  1. Confirm the primary region is really unavailable, not just slow from where you are looking.
  2. Fence the old primary: block its writers at the network or credential level so it cannot accept writes if it returns.
  3. Record the replica's replication position so you know which writes may be lost.
  4. Promote the replica, then point applications at it through configuration or a stable endpoint name.
  5. Shift load balancer capacity to the surviving region and scale it to full peak.
  6. Rebuild the old region as a replica later; never let it rejoin as a primary.

The data tier, service by service

ServiceMulti-region behaviourWhat it means for you
Spanner (multi-region config)Synchronous Paxos replication across regions; 99.999% availability SLA (regional: 99.99%)RPO zero for committed writes; writes pay cross-region quorum latency
Firestore (multi-region location)Read-write replicas in several regions plus a witness; 99.999% SLASurvives a region loss without application failover
Cloud SQLRegional HA across zones; cross-region read replicas are asynchronousEnterprise Plus advanced DR adds replica failover and zero-loss switchover; plan for some loss on unplanned failover
AlloyDBRegional clusters; cross-region secondary clusters replicate asynchronouslyPromotion is a deliberate action with a small RPO
Cloud StorageDual- and multi-region buckets replicate asynchronously: target 99.9% of new objects within 1 hour, 100% within 12 hoursTurbo replication (dual-region only) targets 100% within 15 minutes
BigQuery US / EUData stored in a single region within the multi-regionNo cross-region redundancy unless you configure managed disaster recovery
MemorystoreRegionalTreat caches as per-region and rebuildable

Two rows deserve emphasis. A BigQuery dataset in the US multi-region is not replicated across regions; Google's location documentation says so directly and points to managed disaster recovery for cross-region redundancy. And Pub/Sub Lite, which older designs used for cheap regional streams, was turned down on 18 March 2026; use Pub/Sub or Managed Service for Apache Kafka instead. Cloud Storage details are in GCS regional, dual-region and multi-region.

Spanner placement and read choices

A Spanner multi-region configuration has three replica types. Read-write replicas hold full data and vote. Read-only replicas hold full data, serve low-latency reads and do not vote. Witness replicas vote but do not hold a full copy. In the nam3 configuration, the default leader region us-east4 and the region us-east1 each hold two read-write replicas, and us-central1 holds the witness. Google documents the write quorum as one replica in the default leader region plus any two of the other four voters.

Write latency therefore follows the distance between the two read-write regions, which is why base configurations pair nearby regions. Put the default leader in the region where most writes originate. Configurations that span continents, such as nam-eur-asia1, keep both read-write regions in North America and put read-only replicas elsewhere: reads are local everywhere, writes from Europe or Asia cross an ocean.

Reads are the main lever. A strong read sees every committed write but a replica may need to confirm with the leader that it is current. A stale read at a fixed or bounded age can be served by the nearest replica without that check:

import datetime
from google.cloud import spanner

db = spanner.Client().instance("prod").database("orders")

# Strong read: latest committed data, may contact the leader region.
with db.snapshot() as snap:
    row = list(snap.execute_sql("SELECT status FROM Orders WHERE id=@id",
                                params={"id": order_id},
                                param_types={"id": spanner.param_types.STRING}))

# Exact staleness: data as of 15 s ago, served by the closest replica.
with db.snapshot(exact_staleness=datetime.timedelta(seconds=15)) as snap:
    rows = list(snap.execute_sql("SELECT sku, price FROM Catalog"))

Use strong reads for anything that must reflect a write the user just made, and stale reads for catalogues, feeds and dashboards. Read-write transactions always go through the leader. More on the clock that makes this possible is in the TrueTime article.

The front door: global load balancing

The global external Application Load Balancer announces one anycast IP from Google's edge. A user connects to the nearest edge location, TLS terminates there, and the request travels over Google's network to the closest backend with capacity. When backends in that region are unhealthy or at their configured capacity, traffic goes to the next closest region. There is no DNS change and no TTL to wait for.

for R in us-east4 us-east1; do
  gcloud run deploy api --image=$IMAGE --region=$R
  gcloud compute network-endpoint-groups create api-neg-$R \
      --region=$R --network-endpoint-type=serverless --cloud-run-service=api
done
gcloud compute backend-services create api-bs --global \
    --load-balancing-scheme=EXTERNAL_MANAGED
for R in us-east4 us-east1; do
  gcloud compute backend-services add-backend api-bs --global \
      --network-endpoint-group=api-neg-$R --network-endpoint-group-region=$R
done

Serverless network endpoint groups do not use classic health checks, so a region that is up but returning errors keeps receiving traffic unless you configure outlier detection on the backend service, or route around it yourself. Test this explicitly. With GKE or VM backends, health checks and capacity settings drive failover. The global HTTP(S) load balancer article covers the URL map, backend services and debugging in depth.

Worked example: an order service in two regions

An order service for customers in the eastern United States targets 99.99% availability, RPO zero for orders, and survival of a region outage. Choose Spanner nam3 for orders, Cloud Run in us-east4 and us-east1 behind one global load balancer, Memorystore per region, a dual-region bucket with turbo replication for invoices, and Pub/Sub for order events.

Order writes commit once a quorum in us-east4 and us-east1 acknowledges, one inter-region round trip more than a regional instance. Product pages use 15-second stale reads from the local replica. Each region runs enough instances to carry 100% of peak alone, enforced by Cloud Run minimum instances and maximum instance limits that leave headroom, because failover moves the whole load at once.

Losing us-east1 looks like this: the load balancer stops sending traffic there; Spanner keeps committing with the leader-region replicas plus the witness; us-east4 caches warm up; invoices written in the last few minutes before the outage may be unavailable until replication catches up. Nobody runs a database failover. Losing us-east4 is similar, except that Spanner moves the leader to us-east1 and write latency changes briefly. The plan to rehearse is not a failover script but a capacity and dependency check.

Failure modes

  • The hidden single region. A BigQuery dataset, a regional Cloud SQL instance, a secret or an Artifact Registry repository in one region quietly makes the whole service single-region.
  • No spare capacity. Two regions each at 70% cannot absorb each other. Keep N+1 capacity, and use reservations for GPUs or large machine types that may be scarce during a regional event.
  • Async promotion surprises. Promoting a Cloud SQL or AlloyDB replica loses writes that had not replicated, and the old primary must not accept writes again. Fence it before promoting.
  • Errors that are not outages. A bad deploy in one region returns 500s while health checks pass. Use outlier detection, regional canaries and per-region SLOs.
  • Cross-region egress. Chatty calls from one region to a database leader in another cost latency and money. Keep request paths regional except for the quorum.
  • Global configuration changes. The load balancer, IAM and DNS are global; a bad change affects every region at once. Stage these changes and keep rollback ready.

Operating a multi-region deployment

Monitor availability and latency per region, not just globally, so a degraded region is visible. Track Spanner commit latency and CPU per region and keep CPU under Google's published recommended maximums for multi-region instances. Measure Cloud Storage replication against your RPO. Run a game day each quarter: remove a region from the backend service, scale its Cloud Run service to zero, and watch latency, error rate and the other region's autoscaling. Keep infrastructure as code so a third region can be added in hours. Cloud DNS still matters for names outside the load balancer; see Cloud DNS.

Trade-offs and cost

Multi-region roughly doubles compute and storage cost, adds Spanner multi-region node pricing, and adds cross-region egress. In return, a region outage becomes an event your users barely notice. If your RPO can be minutes and your RTO an hour, a warm standby with asynchronous replication is much cheaper than active-active Spanner. If you need RPO zero with no operator action, Spanner or Firestore multi-region is the simplest correct answer on Google Cloud. Keep the choice per data set: orders on Spanner, analytics on BigQuery with managed DR if needed, images in a dual-region bucket.

What to do next

  1. Inventory every dependency and record its location type: zonal, regional, dual-region or multi-region.
  2. Write RPO and RTO per data set, then choose synchronous or asynchronous replication for each.
  3. Put the Spanner default leader where writes originate and use stale reads where freshness allows.
  4. Front the service with one global external Application Load Balancer and backends in at least two regions.
  5. Configure outlier detection for serverless backends and prove that a region returning errors gets drained.
  6. Provision N+1 capacity and reservations for scarce machine types.
  7. Check BigQuery datasets and replace any Pub/Sub Lite dependency.
  8. Rehearse a region loss every quarter and record the measured RTO and data loss.
Key takeaway: On Google Cloud, compute goes multi-region easily, and the data tier decides everything else. Spanner and Firestore multi-region give RPO zero without operator failover; Cloud SQL, AlloyDB and Cloud Storage replicate asynchronously; BigQuery multi-regions are not replicated at all. Map every dependency, hold spare capacity, drain bad regions automatically, and rehearse.