Server

Thundering herd

Thundering herd is a crowd of waiting requests all rushing at a resource the instant it becomes available. The first request is fine and the rest overwhelm it.

How it is measured

Look for a sharp spike aligned with an event: a cache key expiring, a lock releasing, a server coming back, a cron boundary. Request counts to the origin for one URL jump from near zero to hundreds within a second.

Check cache miss logs for the same key at the same timestamp. If 300 identical misses land inside one second, the herd is there.

Worked example

A homepage cache entry expires at 12:00:00. In the next second 420 visitors miss the cache together, and each triggers the same 1.8 second database aggregation. The database receives 420 copies of one heavy query, CPU hits 100 percent, and pages time out for half a minute.

Request coalescing, so one request rebuilds while others wait, together with serving stale content during refresh, drops the origin load at expiry to a single query.

How it differs

A thundering herd is a simultaneous rush at a freed resource. A retry storm is a crowd created by failed requests being retried. A herd excludes the failure loop, it can happen on a perfectly healthy cache expiry. A retry storm excludes a single trigger moment, and builds over repeated attempts. A cold start can start a herd when many requests hit an unready function.

Common errors

Setting all cache keys to expire at the same minute. Letting every request recompute on a miss. Restarting all servers at once so caches are cold together. Using cron at :00 across the fleet. Treating the spike as an attack.

In practice

Add jitter to cache lifetimes, use stale-while-revalidate or a lock so one request refreshes, and stagger cron jobs and restarts. Warm the cache before a deploy goes live.

See also

Retry storm, Cold start

Sources

Count this on a real site.

Watch my website