Uptime

Active-active

Active-active is a layout where two or more origins serve live traffic at the same time. Losing one halves capacity instead of taking the site down.

How it is measured

Check that every node takes real requests: per-origin request counts from the load balancer, and probe results for each origin address, not just the shared hostname. A healthy setup shows traffic split in a stated ratio, such as 50/50 or 70/30, with both origins passing health checks.

The number that matters is headroom. If each origin runs at 70 percent CPU at peak, losing one pushes the survivor past 100 percent. Measure peak load per node, then multiply by the share it inherits when its twin dies.

Worked example

A WooCommerce shop runs one origin in Frankfurt and one in Virginia behind a DNS load balancer. At 14:10 Frankfurt's database primary stalls. Probes against the Frankfurt IP fail three times in a row, the balancer pulls it, and Virginia takes the full 380 requests per second that it used to handle at half of that.

Virginia's PHP workers saturate at 300 requests per second. For eleven minutes, 80 requests per second return 502s even though one origin is still up. The fix was not a third region. It was sizing each node to carry the whole load.

How it differs

Active-active keeps every origin serving. Active-passive keeps a spare idle until failover. Active-active excludes the cold-start gap but adds work to keep writes and sessions consistent across two live sides. Active-passive excludes that split-write risk but cannot prove the spare works under real load.

Common errors

Sizing each node for half the load. Sharing one database and calling the stack redundant. Health-checking the shared hostname instead of each origin. Letting caches and sessions live on a single node. Never testing by actually switching one side off.

In practice

Pull one origin out of rotation during a quiet hour this week and watch error rate and latency on the other. If the survivor struggles, resize before you add a third node. Write down which component, usually the database, is still a single point of failure.

See also

Active-passive, Multi-region

Sources

Count this on a real site.

Watch my website