Multi-region active-active architecture runs an application in two or more AWS Regions, with each participating Region serving production traffic. It can reduce latency for geographically distributed users and keep some operations available during a regional impairment. Those benefits depend on the application, its data model, and the capacity left after a failure.

The most important distinction is between an active-active application tier and a database that accepts writes in multiple Regions. You can run application servers everywhere while still depending on one database writer. Moving traffic alone will not restore writes if that writer is unavailable.

This guide explains how to choose regional boundaries, route traffic, handle replicated data, and test recovery without assuming that another Region guarantees uninterrupted service.

What active-active means on AWS

In an active-active deployment, users normally reach more than one regional application stack. Each stack should contain the compute, networking, and dependencies needed for the operations it promises to serve. A Region may serve all users, a geographic population, or a set of tenants.

Multi-region architecture complements availability within a Region. A workload still needs appropriate redundancy across Availability Zones (AZs); a second Region does not fix a single-AZ bottleneck in the first.

Common availability and recovery approaches
ApproachNormal operationMain tradeoff
Multi-AZ in one RegionResources are distributed across AZs, according to the service configuration.Protects against some local failures but does not provide another regional application stack.
Multi-region active-passiveOne Region serves traffic; another is prepared for recovery.Ongoing capacity and activation work depend on the standby strategy.
Multi-region active-activeMultiple Regions serve traffic continuously.Requires tested traffic shifts, enough surviving capacity, and explicit cross-region data behavior.

For workloads that can tolerate activation time, pilot light or warm standby disaster recovery may meet the recovery objective with less operating burden. Active-active is useful when normal traffic distribution or recovery requirements justify the extra complexity.

Define the failure and data requirements first

Set a recovery time objective (RTO): the target time to restore the required service. Set a recovery point objective (RPO): the amount of recent data loss the business can tolerate, usually expressed as a time window. These are targets to verify, not properties conferred by a routing service.

Answer these questions for each important user operation:

  • Must it keep working during an AZ failure, a regional outage, or a network partition between Regions?
  • Can it return stale data? Does a user need to read a write immediately after moving to another Region?
  • Can two Regions update the same record? What should happen to conflicting updates?
  • Must several records change atomically, such as a debit and a credit?
  • Which Regions are permitted to store each dataset, including replicas and backups?
  • Can the surviving stack handle the diverted traffic, including database requests and background work?

For example, a replicated product catalog may tolerate a delayed update. Inventory reservations or payment operations need explicit rules for concurrent changes and duplicate requests. The same replication choice need not fit both datasets.

Build regional stacks that can serve independently

A practical starting point is a complete application stack in each chosen Region: a VPC, an ingress endpoint, compute across AZs, and the required regional data and messaging resources. Keep the normal request path regional wherever the consistency requirements allow it. A synchronous call from every Region to one central dependency creates a shared failure point.

AWS's guidance on deploying across multiple locations emphasizes regional independence. VPC peering or Transit Gateway inter-region peering can support specific private connectivity requirements; Direct Connect and Site-to-Site VPN primarily address connectivity with external networks. They are not mandatory components of every active-active application.

Keep durable state outside replaceable compute

Application instances can use local memory or temporary disk for caches and intermediate work. They should not rely on that local state as the only copy of a session, job, or business record. Store durable state in a service whose replication and recovery behavior match the operation.

Externalizing state to one central database does not make a regional stack independent. Session storage, authentication dependencies, queues, secrets, encryption keys, and third-party integrations must also work when the other Region is unavailable. Decide whether a user can continue a session after rerouting and how stale session data is handled.

Deploy compute separately in each Region

Amazon EKS runs Kubernetes; Amazon ECS is a separate container orchestrator. Deploy regional clusters or services for the platform you use. An ECS cluster is Region-specific, and an EKS control plane spans AZs within one Region.

For serverless APIs, deploy the required Lambda functions and regional API Gateway endpoints in each Region, then configure global routing separately. API Gateway does not automatically select a healthy regional copy of your application. Likewise, regional queues and workflows do not become one globally replicated queue or workflow just because the function exists in several Regions.

Choose routing and data services by their actual behavior

Route 53: DNS-based traffic distribution

Route 53 can distribute DNS answers using latency or weighted routing policies and health evaluation. Active-active routing normally includes multiple eligible regional records; the failover routing policy instead describes a primary/secondary arrangement.

Latency routing chooses according to AWS's latency measurements, rather than simply selecting the geographically closest Region. Configure endpoint health checks or supported alias target health evaluation for the resources you use. A load balancer's health status and an application-level ability to complete a transaction are different signals.

DNS changes affect subsequent lookups. Resolvers and clients can retain an earlier answer according to its time to live (TTL), and existing connections do not move to another Region. See the Route 53 FAQ for caching and failover behavior. Shorter TTLs can reduce part of the delay, but they do not remove health detection or application recovery time.

Global Accelerator: a stable network entry point

A standard AWS Global Accelerator provides static IP addresses and routes TCP or UDP traffic over the AWS network to regional endpoints. Its supported endpoint types include Application Load Balancers, Network Load Balancers, EC2 instances, and Elastic IP addresses. An API Gateway endpoint is not a directly supported standard accelerator endpoint.

Global Accelerator can redirect new connections after detecting an unhealthy endpoint. It does not migrate established connections; applications need reconnect and retry behavior. Review how connection routing works and the documented fail-open behavior when no eligible healthy endpoint is found. Health-based routing is not a guarantee that unhealthy endpoints will never receive traffic.

CloudFront caching can reduce latency and origin load for suitable HTTP content. Lambda@Edge can customize requests at the edge, but neither replaces your regional application or resolves database consistency.

DynamoDB global tables: select the consistency mode

DynamoDB global tables support writes in multiple Regions. They offer two consistency modes:

  • Multi-Region eventual consistency (MREC): replication is asynchronous. Concurrent changes to one item use last-writer-wins reconciliation. Even a strongly consistent local read can miss a recent write from another Region. Conditions are evaluated locally, so two Regions can both accept conflicting business actions.
  • Multi-Region strong consistency (MRSC): successful writes are synchronously replicated to at least one other Region; strongly consistent reads return the latest item across replicas. It requires exactly three participating Regions, using three replicas or two replicas and a witness, within a supported Region set. It adds cross-region latency and does not support DynamoDB transaction operations.

The consistency mode cannot be changed after table creation. MREC transactions are atomic only in the originating Region and are not replicated as an atomic unit. Verify supported Regions and features in AWS's global table design guidance before designing around MRSC.

Do not treat either mode as an application-level conflict policy. For MREC, consider assigning each mutable record a home Region or using immutable events where appropriate. Transferring ownership during a failure needs its own safeguards. For MRSC, handle concurrent-write errors and choose the read consistency required by the operation. Neither approach removes the need for safe retries.

Aurora Global Database: one primary writer

Aurora Global Database has one primary write Region and read-only secondary clusters. Replication to secondary Regions is asynchronous. Applications can read locally, but their write path still depends on the primary.

Write forwarding lets supported secondary clusters send write statements to the primary; it does not turn them into independent writers. During a primary-region failure, write recovery requires a supported failover procedure, application reconnection, and consideration of replication lag and potential data loss. A planned switchover and recovery from an unplanned outage are different operations.

Describe this accurately as an active-active application tier with a single-primary relational database. If every Region must continue accepting independent writes during isolation, that requirement needs a different data design.

S3 replication: asynchronous object copies

S3 Cross-Region Replication copies eligible objects asynchronously between buckets. It requires versioning and the appropriate replication permissions. Existing objects need S3 Batch Replication; creating a live replication rule does not backfill them automatically.

If both Regions create objects, configure replication in both directions as needed, with deliberate object-key ownership. Replica modification sync covers supported metadata changes; it is not a general conflict-resolution mechanism. A regional S3 write does not guarantee that another Region can immediately read its replica.

An S3 Multi-Region Access Point can route requests, but does not copy the data by itself. See the S3 cross-region replication use cases for where replicated buckets help, then validate object availability and permissions during traffic shifts.

Implement and validate the complete request path

  1. Build and test one regional stack. Verify its AZ redundancy, capacity limits, dependency behavior, and backups before duplicating it.
  2. Deploy the second stack reproducibly. Use infrastructure as code, with regional configuration for endpoints, keys, quotas, and images. Deploy application changes progressively so a faulty release does not affect every Region at once.
  3. Prepare the data. Configure the chosen replication model, populate initial replicas, check permissions, and measure lag or write latency. Exercise concurrent updates and read-after-write behavior.
  4. Introduce traffic gradually. Check actual latency and errors from the user locations you serve. Confirm that health signals reflect the operations each stack can perform.
  5. Test regional loss. Remove one stack from service in a controlled environment. Include diverted traffic, retries, sessions, and background jobs in the test.
  6. Test the return to service. Reconcile data and backlog, restore healthy capacity, and shift traffic back gradually.

With two Regions normally serving half the traffic each, losing one roughly doubles the survivor's request load. That is a capacity scenario to test, not a reason to assume Auto Scaling will react in time. Preprovision enough headroom for the required recovery behavior, or explicitly define which operations can be degraded or rate-limited.

Retries must be bounded and safe. A timed-out write may already have succeeded, so use idempotency or deduplication for side effects. A deduplication record stored with asynchronous replication can itself lag; account for that before retrying a payment or job in another Region. The cross-service resilience guide explains retries, timeouts, and failure boundaries.

Operate and test the recovery design

Monitor each Region independently and the end-to-end user experience. Track successful requests, errors, latency, saturation, queue age, and the chosen database's replication or consistency signals. For MREC DynamoDB, ReplicationLatency measures propagation between replica pairs; MRSC tables do not publish that metric.

A useful recovery exercise tests more than a traffic switch:

  • Can the surviving Region accept the required reads and writes under the new load?
  • What happens to existing connections and requests whose outcome is unknown?
  • Do asynchronous jobs disappear, remain in the impaired Region, or run twice?
  • Can operators recover without first deploying resources or changing configuration in the impaired Region?
  • Does failback preserve acknowledged data and avoid replaying completed work?

Record measured recovery time and the actual data recovery point. Controlled fault injection and disaster recovery testing answer different questions; use the test that exercises your runbook and its dependencies.

Keep restorable backups as well as live replicas. Replication can propagate unwanted changes, and a healthy replica is not necessarily a clean recovery point. AWS's disaster recovery guidance explains why multi-site operation still needs data protection and tested recovery.

Budget for duplicate infrastructure, replicated writes, storage, inter-region transfer, monitoring, and spare failover capacity. Active-active can improve regional availability and global response times, but its value comes from meeting a tested workload requirement, not from minimizing the number of running instances.

Frequently asked questions

Does active-active guarantee zero downtime or zero data loss?

No. Routing detection, cached DNS, connection recovery, capacity shortages, and application dependencies can interrupt service. Data loss depends on the storage and replication model. MRSC DynamoDB supports zero RPO for successful writes, but that does not guarantee uninterrupted operation for the complete application.

Do all Regions need private network connectivity?

Only when a dependency requires it. Some managed replication and public API paths do not need VPC-to-VPC connectivity. Define the necessary data flows and security controls rather than adding a shared network dependency by default.

When is active-active worth the complexity?

When serving users from multiple Regions or meeting a regional recovery objective justifies the additional data, capacity, deployment, and operating work. Compare it with a tested multi-AZ or active-passive design against the same requirements.