Uptime

MTTR

Also called mean time to restore.

MTTR is the average time from the start of a failure to restored service. Teams use it to judge how fast they recover once something breaks.

How it is measured

For each incident, subtract failure start from restore time. Average across the period. Define which R you mean (restore, repair, recover, or respond) and keep it constant. Look at median and worst too, because one long outage dominates the mean.

Break it into detect, acknowledge, diagnose, fix, and verify. The largest slice shows where to spend effort.

Worked example

A travel booking site had four incidents in a quarter, lasting 12, 18, 25, and 140 minutes. MTTR is 48.75 minutes. The median is 21.5. The 140-minute incident was a corrupted search index that needed a full rebuild.

Keeping a hot standby index brings a repeat of that event to about 20 minutes. The mean would fall to about 19 minutes with no change to the other three.

How it differs

MTTR spans failure to restored service. MTTA spans alert to acknowledgement only. MTTR includes diagnosis and the fix. MTTA excludes them.

Common errors

Averaging long and short events with no median. Changing the definition of resolved. Starting at the alert instead of the failure. Leaving reopened incidents out. Closing tickets early to improve the number.

In practice

Compute last quarter's MTTR and median from incident records. Break the longest event into stages. Choose one stage, such as rollback speed or runbook quality, and cut it.

See also

MTTA, RTO

Sources

Count this on a real site.

Watch my website