Uptime

Alert fatigue

Alert fatigue is what happens when people get so many notifications that they stop reading them. The page that matters gets swiped away with the forty that did not.

How it is measured

It is not a metric a probe can give you, so observe it. Count pages per on-call shift, the share that needed action, and the median time to acknowledge. A common warning sign is fewer than half of pages leading to any change.

Also watch behavior: muted channels, auto-acknowledge rules, and mail filters that send alert email to a folder. Those habits prove fatigue before any survey does.

Worked example

A two-person agency monitors 60 client WordPress sites with a 2-minute HTTP check and no confirmation step. Over one week the on-call phone gets 212 pages. 190 clear on their own within five minutes, mostly slow shared hosts. 9 are real outages.

On Thursday night a client's checkout is down for 40 minutes because the page sat unread at 02:15 among three that had cleared themselves. After the agency requires two failed probes from two regions and moves slow-response warnings to a morning digest, weekly pages fall to 14, and 8 of those need action.

How it differs

Alert fatigue is the human outcome. Flapping is one of its biggest causes: a single check toggling up and down, each flip a new page. Flapping is a property of a signal. Fatigue is a property of the people receiving many signals, and it can exist with no flapping at all.

Common errors

Paging on warnings. Alerting on every single failed probe. Sending the same alert to five channels. Never deleting alerts nobody acts on. Treating more alerts as more safety.

In practice

Export last month's pages and mark each one actionable or not. Delete or demote any alert that was not actionable three times. Keep pages for symptoms users feel and send the rest to a ticket or a daily digest.

See also

Flapping, On-call

Sources

Count this on a real site.

Watch my website