Uptime
Flapping
Flapping is a check or alert that flips between up and down repeatedly in a short time. Each flip can send a new notification.
How it is measured
Count state transitions in a window. More than about 4 changes in 30 minutes marks a flapping check. Plot probe results as a strip: red, green, red, green stands out at a glance.
Find where it flips: one region, one node behind a balancer, or one timeout sitting near its limit. A response time hovering around a 5-second limit flaps. A hard failure does not.
Worked example
A forum on shared hosting has a monitor with a 3-second timeout. Page generation takes 2.7 to 3.4 seconds at night. Between 01:00 and 02:00 the monitor logs 19 transitions and the on-call phone buzzes 19 times, while readers saw a working site.
Raising the timeout to 8 seconds and requiring 3 consecutive failures cuts that to a single alert in a week, a real 6-minute outage.
How it differs
Flapping is one signal toggling. Alert fatigue is the team-wide result of many noisy signals. Flapping excludes any claim about people and points to a technical cause. Fatigue excludes the cause and points to the load on humans.
Common errors
Alerting on every transition. Timeouts set at the median response. Probing from a single region. Treating each flap as a new incident. Requiring no successes in a row before declaring recovery.
In practice
Find the five checks with the most transitions last week. Raise confirmations to 3, add a second region, and require 2 successes to clear. Fix the underlying slow endpoint if that is the cause.