Uptime

RTO

Also called recovery time objective.

RTO is the longest downtime you will accept after a failure before service must be back. It is a target for recovery time.

How it is measured

Test against the real clock: from declaring the failure to verified working, including spinning up, restoring data, DNS, and smoke tests. A drill gives a number. A design document gives a hope.

Break it into steps with owners, each with a time budget. If restore alone takes 90 minutes, an RTO of 30 is fiction.

Worked example

A legal-documents portal sets an RTO of 60 minutes. A drill: provision a VM (8 minutes), install the stack from a script (14), restore a 22 GB database (31), update DNS (5), run smoke tests (6), and wait out a 5-minute cache on the old record.

The drill takes 69 minutes, missing the target by nine. They pre-bake an image and keep a nightly-restored warm database, cutting the drill to 24 minutes.

How it differs

RTO is time to be back. RPO is how much data is lost. RTO excludes data age. RPO excludes how long you wait for service.

Common errors

Picking a number without a drill. Counting only the technical restore and not decision time. Forgetting the DNS TTL. Assuming the one engineer who knows how is available. Using the same RTO for every system.

In practice

Run a restore drill and time each step. Compare the total to your RTO. Shorten the slowest step first, usually data restore or DNS.

See also

RPO, MTTR

Sources

Count this on a real site.

Watch my website