Uptime

Escalation

Escalation is the path an alert follows when the first responder does not respond in time or needs help. It names who gets paged next and how long to wait.

How it is measured

Define tiers with timers: primary on-call at 0 minutes, secondary after 10, engineering manager after 25. Measure the share of pages that needed tier two and the time between tiers.

Test by running an unacknowledged alert through the whole chain in a drill. Check phone numbers, override settings, and that each tier knows it is next.

Worked example

A two-store e-commerce company pages Priya at 02:40 about a failing checkout. She is on a flight with no signal. After 10 minutes with no ack, the tool pages Mateo, whose phone is on silent with a do-not-disturb rule blocking the call. At minute 25 the owner's phone rings and he starts the restore.

Detection to first human action took 27 minutes. The postmortem fixes two things: critical pages bypass silent mode, and a secondary is added whenever the primary is marked traveling.

How it differs

Escalation is moving up the chain when no one answers. Paging is the single act of alerting a person. Paging excludes who comes next. Escalation excludes the earlier decision of what deserves a page in the first place.

Common errors

A chain with only one person. Ten-minute waits for a total outage. Out-of-date phone numbers. No path for the responder who needs help. Never testing the third tier.

In practice

Draw your chain on one page with names, times, and channels. Run an unacknowledged test page after hours this month. Remove people who left, and add a secondary for holidays.

See also

Paging, On-call

Sources

Count this on a real site.

Watch my website