Uptime

Paging

Paging is the act of sending an alert that is allowed to interrupt a person right now, through a phone call, SMS, or push that bypasses silent mode.

How it is measured

Measure delivery and response: percent of pages delivered within 30 seconds, ack time, and pages per week per person. Include a test page in each shift handover.

Gate pages by criteria: user-facing impact, needs a human now, and the action is known. Anything else is a ticket or a chat message.

Worked example

A logistics dashboard routes 41 alerts a week to paging. Auditing a month shows 29 were CPU warnings on a batch box that recovered alone, 8 were disk-at-80-percent warnings, and 4 were real customer-facing 5xx runs.

After the audit, only the 5xx and checkout probes can page. Weekly pages fall to 5. The disk warning becomes a ticket with a daily summary.

How it differs

Paging is the interruption. On-call is the commitment to respond. Paging excludes the schedule and ownership. On-call excludes the delivery channel and its settings.

Common errors

Paging on causes instead of symptoms. SMS-only with no phone call. A silent-mode bypass that was never set up. The same page going to the whole team. No test page. Pages without a runbook link.

In practice

Review each paging rule against three tests: users are hurt, a human is needed now, and the first step is known. Attach a runbook link in the page body. Send a test page on the first day of each rotation.

See also

On-call, Escalation

Sources

Count this on a real site.

Watch my website