Uptime

On-call

On-call is the rotation that makes one named person responsible for responding to alerts outside normal hours. It trades their free time for response speed.

How it is measured

Track pages per shift, night pages, time to acknowledge, and the share of shifts that lose a night's sleep. Note rotation length (one week is common) and team size (at least four or five people for a sustainable load).

Check coverage: every hour has a primary, a secondary, and a handover. A calendar gap is an outage waiting.

Worked example

A six-person team covering a SaaS runs week-long shifts. In March, Mei's week brings 3 night pages. Jon's has 17, from a flapping disk check on a staging box nobody turned off. Jon works the next day on 3 hours of sleep and ships a bug.

The team removes the staging page, sets a cap of 5 night pages per shift before the load is reviewed, and adds a stipend plus time off after bad weeks.

How it differs

On-call is the duty roster. Paging is the notification that wakes the person. On-call excludes the tool and the message. Paging excludes who is responsible and for how long.

Common errors

A one-person rotation. No handoff notes. Paging for non-urgent issues. No compensation. No secondary. No shadow shift for new people.

In practice

Write down the rota for the next eight weeks, with a secondary. Check each person has access, runbooks, and a test page before their shift. Review last month's night pages and kill the noisy ones.

See also

Paging, Escalation

Sources

Count this on a real site.

Watch my website