Uptime

Multi-region

Multi-region is running a service in more than one geographic region so a failure in one does not take the whole site down.

How it is measured

Verify each region with its own probes and compare latency per region. Record where data lives, which region takes writes, and how long replication lags between them.

Test region loss by blocking one in staging. Count what breaks: sessions, caches, queued jobs, and third-party callbacks hardcoded to one hostname.

Worked example

A video-course platform serves from us-east and eu-west, with the database primary in us-east and a read replica in eu-west about 400 ms behind. When us-east loses a zone at 13:00, reads continue from eu-west, but logins need to write a session row and fail for 9 minutes until the replica is promoted.

The site was multi-region for static pages and single-region for anything with a write. Promoting automatically after 60 seconds, with sessions moved to a signed cookie, closes most of that gap.

How it differs

Multi-region spreads a service across geographies. Failover is the switch from a failed place to a working one. Multi-region without failover logic is two separate sites. Failover without multi-region is a move between machines in one room.

Common errors

A single write database. Sharing one DNS provider. Using the same cloud account for everything. Cross-region latency surprises. No test of region loss. Overlooking data residency rules.

In practice

List, per feature, which region it needs for a write, and name the single point that remains. Run one region-loss game day a year. If the budget is not there, a cold copy and a tested backup is the sensible first step.

See also

Failover, Probe

Sources

Count this on a real site.

Watch my website