An AWS Region is a fault domain. Availability Zones protect you from a data centre losing power; they do not protect you from a regional service impairment, a bad deployment of your own that reaches every zone, or a requirement to keep serving if a whole Region is unreachable. Multi-Region architecture is the answer to those cases, and it is also one of the most expensive and error-prone designs on the platform. Most teams that run it discover that the hard part is not starting a second copy of the stack but deciding where writes go, how much data a failover may lose, and who presses the button.

This article builds the design from two numbers, the recovery point objective (RPO, how much recent data you may lose) and the recovery time objective (RTO, how long you may be down), then walks through routing, data replication, a failover runbook, a worked example and the failure modes that make real failovers go wrong.

Warm-standby multi-Region: routing, regional stacks and replicated stateUsersresolve api.example.comRoute 53health checks, failover recordsARC routing controls5-Region cluster, On/OffDNShealth check stateOperatordecides failoverRegion A: us-east-1 (primary, takes writes)ALB + servicesECS or EKS, scaledQueues + cacheSQS, ElastiCacheAurora primarywriter clusterDynamoDB replicaglobal tableS3 bucket + KMS multi-Region keyreplication rules, Secrets replicasRegion B: us-west-2 (warm standby)ALB + servicesrunning, scaled downQueues + cacheregional, not sharedAurora secondaryread-only until promotedDynamoDB replicaaccepts writes (LWW)S3 replica bucket + same key IDstorage-level async replicationS3 replicationprimary answersecondary answerFailover = fence writes in A, promote Aurora in B, flip routing controls, scale B up.Every step uses data-plane operations, so it works while A's control plane is impaired.
Warm standby across two Regions: Route 53 health checks backed by ARC routing controls choose the Region; data services replicate asynchronously.

Choose a strategy from RPO and RTO

AWS's disaster recovery guidance names four strategies. They differ in what runs in the second Region before anything goes wrong, and that single choice sets both cost and RTO.

StrategyWhat runs in Region BTypical RTORPO driver
Backup and restoreNothing; backups are copied thereHoursBackup frequency
Pilot lightData replicas only; compute defined in code but offTens of minutesReplication lag
Warm standbyFull stack at reduced capacityMinutesReplication lag
Active-activeFull stack serving live trafficNear zero for readsReplication lag and conflict rules

Cost rises down the table, because you pay for idle or duplicated capacity, cross-Region data transfer and the engineering time to keep two Regions identical. The RPO column is the important one: every strategy except synchronous replication loses whatever had not yet replicated when the primary failed. Pick the cheapest row whose RTO and RPO your business has actually agreed to, in writing.

Choose a strategy from RPO and RTO

AWS's disaster recovery guidance names four strategies. They differ in what runs in the second Region before anything goes wrong, and that single choice sets both cost and RTO.

StrategyWhat runs in Region BTypical RTORPO driver
Backup and restoreNothing; backups are copied thereHoursBackup frequency
Pilot lightData replicas only; compute defined in code but offTens of minutesReplication lag
Warm standbyFull stack at reduced capacityMinutesReplication lag
Active-activeFull stack serving live trafficNear zero for readsReplication lag and conflict rules

Cost rises down the table, because you pay for idle or duplicated capacity, cross-Region data transfer and the engineering time to keep two Regions identical. The RPO column is the important one: every strategy except synchronous replication loses whatever had not yet replicated when the primary failed. Pick the cheapest row whose RTO and RPO your business has actually agreed to, in writing.

Routing traffic and failing over on the data plane

Traffic reaches a Region in one of two ways. Route 53 answers DNS queries using routing policies such as failover, latency, weighted or geolocation, and attaches health checks so an unhealthy endpoint stops being returned. Failover through DNS takes at least the record's TTL, usually 60 seconds or less for these records, plus however long resolvers and clients cache answers beyond that TTL; some clients cache far longer than they should. Global Accelerator gives you static anycast IP addresses and shifts traffic between Regional endpoints without depending on client DNS caches, at extra cost.

The design rule that matters most is to fail over using the data plane, not the control plane. Route 53 answering queries and evaluating health checks is data plane and is built for very high availability. Changing a record through the Route 53 API is a control-plane operation, and control planes are exactly what tends to be impaired during a large event. So pre-create both records and let a health check decide which one is answered.

Amazon Application Recovery Controller (ARC, formerly Route 53 ARC) makes that health check something you control. A routing control is an On/Off switch that backs a Route 53 health check. Its state lives in a cluster of five redundant Regional endpoints (us-east-1, us-west-2, eu-west-1, ap-northeast-1 and ap-southeast-2); ARC keeps at least three of the five reachable for state changes, and AWS documents that state converges across the endpoints in about 5 seconds on average and no more than 15 seconds. Clients must retry across all five endpoints:

import random, boto3

# Cluster endpoints come from describe-cluster; store them with your runbook,
# because you must not depend on a control-plane call during the event.
ENDPOINTS = {
    "us-east-1": "https://aaaa.route53-recovery-cluster.us-east-1.amazonaws.com/v1",
    "us-west-2": "https://bbbb.route53-recovery-cluster.us-west-2.amazonaws.com/v1",
    # ... the other three
}

def set_control(arn, state):
    for region, url in random.sample(list(ENDPOINTS.items()), len(ENDPOINTS)):
        try:
            client = boto3.client("route53-recovery-cluster",
                                  region_name=region, endpoint_url=url)
            client.update_routing_control_state(
                RoutingControlArn=arn, RoutingControlState=state)
            return region
        except Exception as exc:          # try the next endpoint
            print(f"{region} failed: {exc}")
    raise RuntimeError("no ARC endpoint accepted the change")

set_control(PRIMARY_CELL_ARN, "Off")
set_control(STANDBY_CELL_ARN, "On")

Safety rules on the control panel, for example 'at least one cell must be On', stop a tired operator from switching everything off.

Replicating state

Compute is stateless and easy to duplicate. State decides the architecture, and each AWS data service replicates with different semantics.

Aurora Global Database. One primary cluster takes writes; up to several secondary clusters in other Regions receive storage-level asynchronous replication, typically with sub-second lag but with no hard bound. A switchover (previously called managed planned failover) synchronises the secondary before promoting it, so RPO is zero, but it requires a healthy primary and is for planned moves. A failover promotes a secondary during an outage, and AWS states its RPO is typically seconds and depends on the replication lag at the moment of failure. Watch AuroraGlobalDBReplicationLag continuously, because that number is your real RPO.

DynamoDB global tables. The default multi-Region eventual consistency mode makes every replica writable and resolves concurrent writes to the same item with last writer wins. That is simple and dangerous for counters, balances and inventory, where two Regions updating the same item silently drops one update. Multi-Region strong consistency (MRSC) instead gives strongly consistent reads and zero RPO, at the cost of higher write latency and a strict topology: exactly three Regions within one region set, either three replicas or two replicas plus a witness.

S3. Cross-Region Replication copies new objects asynchronously; Replication Time Control adds a 15-minute objective with metrics. Existing objects need Batch Replication. Supporting services must exist in both Regions too: KMS multi-Region keys share key material and key ID so replicated ciphertext decrypts in Region B, Secrets Manager can replicate secrets, and ElastiCache global datastores replicate a cache if you cannot simply let Region B's cache start cold.

A failover runbook

A failover is a sequence, and the order matters because the goal is never to have two writers.

  1. Detect. Alarms on customer-facing error rate and latency from canaries in other Regions, not only on your own health endpoints inside the failing Region.
  2. Decide. A human makes the call against a written threshold. Fully automatic Regional failover tends to trigger on partial impairments and on monitoring failures, and a false failover is itself an outage with data loss.
  3. Fence. Stop writes in Region A: set its routing control Off and, if you can, flip an application-level read-only flag held in Region B.
  4. Promote. Fail over the Aurora global database to Region B and record the replication lag at that moment, which is the data you may have lost.
  5. Route. Turn Region B's routing control On.
  6. Scale. Raise Region B's capacity to full; the warm standby was running at a fraction.
  7. Verify and reconcile. Run synthetic transactions, then reconcile writes made in A after the last replicated point, using the old cluster once it is reachable.

Failback is the same procedure in reverse, but with a healthy original Region you can use a switchover and lose nothing.

Worked example: an order service

An e-commerce order service agrees an RPO of 5 seconds and an RTO of 15 minutes. That rules out backup and restore and pilot light (compute takes too long to appear at scale) and does not justify active-active writes. The design is warm standby across us-east-1 and us-west-2.

  • Orders and payments live in Aurora Global Database, because they need relational constraints and a single writer. A lag alarm at 2 seconds gives early warning against the 5-second RPO.
  • Shopping carts live in a DynamoDB global table in default mode, served locally in both Regions; last writer wins is acceptable for a cart.
  • Every payment request carries an idempotency key stored with the order. After a failover, clients retry, and the key prevents a charge that was captured but not replicated from being taken twice once reconciliation compares the payment provider's records with the database.
  • Region B runs the full service at 20 percent of production capacity, with EC2 or Fargate quotas already raised to 100 percent, because quota increases are a control-plane request that can be slow during an event.
  • Every quarter the team runs a game day: switch over during business hours, serve from Region B for a day, switch back, and record the measured RTO.

The rough cost is a second, smaller stack plus Aurora replica instances and cross-Region transfer, which is why the recommendations and reporting services, which could tolerate an hour of downtime, stayed single-Region with backups.

Failure modes

  • Split-brain writes. Region A is degraded but alive, Region B is promoted, and both accept writes for minutes. Fence before you promote.
  • Hidden single-Region dependencies. A third-party API, a license server, a CI pipeline or a global-service control plane that exists only in one Region. Find them by actually serving from Region B.
  • Configuration drift. The standby's task definitions, AMIs, feature flags or IAM policies fall behind. Deploy both Regions from the same pipeline in every release.
  • Capacity on demand. Scaling Region B from 20 to 100 percent while everyone else is doing the same. Reserve capacity for the critical tier.
  • Flapping health checks. Aggressive thresholds move traffic back and forth. Require several consecutive failures and keep the final decision with ARC and a human.
  • Last-writer-wins surprises. Concurrent updates to the same DynamoDB item from two Regions lose one of them without an error.

Operating and observing two Regions

A multi-Region design is only as good as the evidence that the standby works today. Three habits supply that evidence.

Measure from outside. Run synthetic canaries, for example CloudWatch Synthetics, from Region B against Region A and from Region A against Region B, so an impaired Region's own monitoring is never the only witness. Alarm on customer-facing symptoms: error rate, p99 latency and checkout success.

Treat lag as a service-level indicator. Put AuroraGlobalDBReplicationLag, DynamoDB's ReplicationLatency and S3 replication metrics on the same dashboard as availability, with alarms set well below the agreed RPO. A lag that creeps up every evening during batch loads is a promise you are breaking without noticing.

Prove parity continuously. Deploy both Regions from one pipeline, run the same smoke tests against both after every release, and diff infrastructure state between them on a schedule. A small percentage of real traffic routed to Region B through a weighted record keeps its caches, connection pools and dependencies exercised, and turns a dormant standby into one you have evidence for.

Trade-offs

Active-active buys near-zero RTO for reads and spreads latency benefits across users, but forces you to design every write path for conflicts or to partition users by home Region. Warm standby keeps a single writer and simple consistency, and pays for it with minutes of RTO and seconds of RPO. Strong cross-Region consistency (DynamoDB MRSC, or Aurora DSQL if it fits your workload) removes data loss at the price of write latency bounded by inter-Region round trips. Many organisations get most of the benefit by making only the few revenue-critical services multi-Region and leaving the rest single-Region with tested backups.

What to do next

Related reading: Route 53 routing and health checks, DynamoDB global tables, Aurora internals, S3 replication, disaster recovery planning and chaos engineering. Then:

  1. Get written RPO and RTO targets per service and choose a strategy per service, not per company.
  2. List every stateful store and record its replication mode, lag metric and conflict rule.
  3. Pre-create DNS records and health checks for both Regions; move failover onto ARC routing controls.
  4. Raise service quotas in the standby Region to full production levels now.
  5. Write the fence, promote, route, scale, verify runbook and store ARC endpoints with it.
  6. Run a switchover game day and publish the measured RTO and RPO.

Key takeaway: Multi-Region design starts from agreed RPO and RTO, not from a diagram. Keep a single writer unless you have designed for conflicts, fail over with data-plane mechanisms such as health checks and ARC routing controls, treat replication lag as your live RPO, and prove the whole thing with regular switchover game days.