Uptime

Heartbeat

Heartbeat is a periodic signal a process sends to show it is alive. The monitor alarms when the signal stops.

How it is measured

The worker pushes a ping every N seconds to a monitor. The monitor expects one within N plus a grace of maybe 50 percent. Missing two consecutive beats trips the alarm.

Because the direction is reversed (inbound, not polled), it works behind firewalls and for processes with no URL. Include a payload, like queue depth, so the beat carries state and not only liveness.

Worked example

A queue worker on a small VPS processes order emails and pings every 60 seconds. On Wednesday at 16:20 it blocks on an SMTP call with no timeout. The process still exists, but beats stop. At 16:23, after 3 missed beats, the alarm fires.

An is-the-process-running check would have stayed green. 214 order emails waited until the worker was restarted at 16:31.

How it differs

Heartbeat is a steady signal from inside a process. Cron monitor adds a schedule for batch jobs. A heartbeat has no concept of runs at 02:00 and expects a constant rhythm. A cron monitor expects nothing between runs.

Common errors

Sending the beat from a separate thread that keeps running when the work is stuck. Setting grace below the normal jitter. Using one heartbeat for many workers. No alert when beats resume. Beating before the work is done.

In practice

Put the beat inside the work loop after a real unit of work completes, not on a timer thread. Include a counter or queue depth. Set grace to at least twice the longest normal gap.

See also

Cron monitor, HTTP check

Sources

Count this on a real site.

Watch my website