Uptime

Runbook

Runbook is a step-by-step page an on-call person follows when a particular alert fires. It turns a 3 a.m. problem into a checklist.

How it is measured

It is not a metric, so check it by use. Each paging alert should link to one runbook that opens with symptom, impact, first checks, fixes, and who to escalate to. Measure freshness: last-edited date and the last time someone ran it.

Test by handing it to someone who did not write it. If they cannot complete step three without asking, the runbook is a note.

Worked example

A disk-above-90-percent alert on a log server fires at 04:12. The runbook says: 1) run df -h and du -sh /var/log/* to find the biggest, 2) rotate nginx logs with logrotate -f, 3) if still above 85, delete journal older than 7 days, 4) if the culprit is mysql binlog, stop and page the DBA.

A new hire follows it, reaches step 4 because binlogs are the culprit, stops as instructed, and pages the DBA instead of deleting files. The fix takes 9 minutes with no data lost.

How it differs

Runbook tells you what to do during an incident. Postmortem explains what happened after. A runbook excludes cause analysis. A postmortem excludes live steps. Postmortems should produce runbook edits.

Common errors

Written once, never updated. Steps that say investigate. Missing commands. Hosted on the site that is down. Being too long. No expected output.

In practice

For each of your five noisiest paging alerts, write a runbook with commands and expected output. Host them somewhere that survives your outage. After each page, fix one line.

See also

Postmortem, On-call

Sources

Count this on a real site.

Watch my website