Uptime

Cron monitor

Also called dead man's snitch.

Cron monitor is a check that expects a scheduled job to phone home and raises an alarm when it does not. The failure it catches is silence.

How it is measured

Each job pings a unique URL on success, optionally with start and finish pings. The monitor knows the schedule and a grace period, say nightly at 02:00 with 15 minutes of grace, and alerts if no ping lands by 02:15.

Record duration too. A backup that used to run 6 minutes and now runs 70 is failing slowly. Send explicit failure pings with the exit code so a crashed job does not wait for the grace timer.

Worked example

A WordPress multisite runs a nightly search reindex via system cron at 01:30. The script ends with a curl to a monitor URL. On the 12th a PHP upgrade moves the CLI binary and cron fails silently. No ping arrives by 01:50.

The alert fires at 01:50. Without it, the stale index would have surfaced on the 15th when a client asked why new products were not searchable. Three days of missing search results is the cost avoided.

How it differs

Cron monitor watches a job by its missed signal on a known schedule. Heartbeat is the general version: any process sending a periodic ping. A cron monitor adds a schedule and a grace window. A heartbeat just expects regular beats and has no concept of runs at 02:00.

Common errors

Pinging at the start only, so crashes pass. Grace set too tight, so slow runs page. Pinging even when the work failed. Sharing one URL across many jobs. Relying on WP-Cron, which only runs when a visitor arrives. Forgetting to re-add the monitor after migrating servers.

In practice

List every scheduled job on the site: backups, renewals, imports, email digests. Give each one a monitor URL and a grace period matching its slowest normal run. Ping only after the last command succeeds.

See also

Heartbeat, WP-Cron

Sources

Count this on a real site.

Watch my website