Server

Liveness probe

Liveness probe is a check the platform runs to decide whether a process is stuck and needs to be killed and restarted. It is not a test of whether the app can take traffic.

How it is measured

Configure the path or command, the period (say 10 seconds), the timeout, and the failure threshold (say 3). On Kubernetes these are `periodSeconds`, `timeoutSeconds` and `failureThreshold`. The container restarts after the threshold is reached.

Watch the restart count per pod. A count that climbs by one every few hours means the probe is catching real hangs, and a count that climbs during every traffic spike means the probe is too strict.

Worked example

A Python worker occasionally deadlocks on a lock inside a library, answering nothing but staying alive. Its liveness probe hits /alive, which only reads an in-memory heartbeat updated every 5 seconds. After three missed beats, 15 seconds, Kubernetes restarts the pod.

Previously the deadlocked pod sat idle for hours until someone noticed the queue length. The probe turned a six-hour stall into a 15-second one.

How it differs

A liveness probe asks if the process should be restarted. A readiness probe asks if it should receive traffic now. Liveness excludes temporary busy states, and failing it kills the container. Readiness excludes the restart and only removes the pod from the service. Using liveness for a slow dependency can cause a restart loop.

Common errors

Checking the database in the liveness probe so an outage restarts every pod. Using a timeout shorter than a normal GC pause. Setting the initial delay below the boot time. Making liveness and readiness the same endpoint. Treating restarts as free.

In practice

Keep liveness cheap and local: is the main loop running? Put dependency checks in readiness. Set the initial delay above your slowest boot, and review restart counts weekly.

See also

Readiness probe, Health check

Sources

Count this on a real site.

Watch my website