Uptime
RTO
Also called recovery time objective.
RTO is the longest downtime you will accept after a failure before service must be back. It is a target for recovery time.
How it is measured
Test against the real clock: from declaring the failure to verified working, including spinning up, restoring data, DNS, and smoke tests. A drill gives a number. A design document gives a hope.
Break it into steps with owners, each with a time budget. If restore alone takes 90 minutes, an RTO of 30 is fiction.
Worked example
A legal-documents portal sets an RTO of 60 minutes. A drill: provision a VM (8 minutes), install the stack from a script (14), restore a 22 GB database (31), update DNS (5), run smoke tests (6), and wait out a 5-minute cache on the old record.
The drill takes 69 minutes, missing the target by nine. They pre-bake an image and keep a nightly-restored warm database, cutting the drill to 24 minutes.
How it differs
RTO is time to be back. RPO is how much data is lost. RTO excludes data age. RPO excludes how long you wait for service.
Common errors
Picking a number without a drill. Counting only the technical restore and not decision time. Forgetting the DNS TTL. Assuming the one engineer who knows how is available. Using the same RTO for every system.
In practice
Run a restore drill and time each step. Compare the total to your RTO. Shorten the slowest step first, usually data restore or DNS.