Uptime
Runbook
Runbook is a step-by-step page an on-call person follows when a particular alert fires. It turns a 3 a.m. problem into a checklist.
How it is measured
It is not a metric, so check it by use. Each paging alert should link to one runbook that opens with symptom, impact, first checks, fixes, and who to escalate to. Measure freshness: last-edited date and the last time someone ran it.
Test by handing it to someone who did not write it. If they cannot complete step three without asking, the runbook is a note.
Worked example
A disk-above-90-percent alert on a log server fires at 04:12. The runbook says: 1) run df -h and du -sh /var/log/* to find the biggest, 2) rotate nginx logs with logrotate -f, 3) if still above 85, delete journal older than 7 days, 4) if the culprit is mysql binlog, stop and page the DBA.
A new hire follows it, reaches step 4 because binlogs are the culprit, stops as instructed, and pages the DBA instead of deleting files. The fix takes 9 minutes with no data lost.
How it differs
Runbook tells you what to do during an incident. Postmortem explains what happened after. A runbook excludes cause analysis. A postmortem excludes live steps. Postmortems should produce runbook edits.
Common errors
Written once, never updated. Steps that say investigate. Missing commands. Hosted on the site that is down. Being too long. No expected output.
In practice
For each of your five noisiest paging alerts, write a runbook with commands and expected output. Host them somewhere that survives your outage. After each page, fix one line.