Server
Retry storm
Also called thundering herd.
Retry storm is what happens when many clients retry failed requests at the same time, multiplying load on an origin that is already struggling. The retries cause the outage to last longer.
How it is measured
Look for request rate rising while the success rate falls. If 1,000 requests per second normally arrive and 3,000 are arriving during a failure, with 2,000 being retries, you are in one. Log an attempt number or a retry header so the two can be separated.
Retries per original request is the key ratio. With 3 retries and no backoff, a 100 percent failure rate quadruples load.
Worked example
A mobile app calls an order API. A database failover makes the API return 500 for 40 seconds. The app retries three times, 1 second apart, so the 800 requests per second on the origin become 3,200. The recovered database is flooded and falls over again at second 55.
Using exponential backoff with jitter, and a cap of 2 retries, held the load to about 1,300 during the next failover and the origin recovered in 12 seconds.
How it differs
A retry storm is retries piling on a failing service. A circuit breaker stops the calls on the caller side. The storm excludes any limiting behaviour, it is what happens without one. The breaker excludes the extra traffic, it is the control that prevents the storm. Thundering herd is related but is a crowd arriving at the same instant, not through retries.
Common errors
Retrying immediately with no delay. Retrying at every layer, so a client, a gateway and a service each multiply it. Retrying non-idempotent calls and double-charging. Using the same interval for every client so they stay in step. Not capping retries.
In practice
Audit your clients and SDKs for retry settings. Use exponential backoff with random jitter, limit attempts, and retry only idempotent requests. Add a circuit breaker where a failure would flood a shared dependency.