Uptime

Brownout

Also called degraded.

Brownout is when a site still answers but is slow or partly broken enough that users feel an outage. It sits between fully up and fully down.

How it is measured

Define it with thresholds a probe can test: response over 4 seconds, 5xx above 2 percent of requests, or a key transaction such as search or cart failing while the homepage returns 200. A brownout is invisible to a plain up-or-down check.

Use percentile latency over short windows, not averages. A p95 of 9 seconds with a p50 of 600 ms is a brownout for one user in twenty, which an average of 1.2 seconds hides.

Worked example

A ticketing site on a single MySQL host sells 2,000 seats at 10:00 on Friday. At 10:03 the database hits 100 percent of its connections. The homepage, cached by the CDN, returns 200 in 90 ms. /checkout takes 14 seconds and 31 percent of attempts time out.

The uptime tool reports 100 percent for the hour. Support gets 140 emails. The incident log later records 22 minutes of degradation, not zero.

How it differs

Brownout is slow or partial failure. Outage is no useful answer at all. A brownout excludes a clean down signal, so a plain ping stays green while users leave. An outage excludes the ambiguity: nothing works.

Common errors

Only probing the cached homepage. Using average latency. Calling it up because the status is 200. Having no threshold, so each person decides alone. Leaving brownout minutes out of the availability math.

In practice

Add one probe on a path that hits the database and one latency threshold, and page when it fails twice in a row. Decide in writing whether brownout minutes count against your SLO. Label the event degraded on the status page rather than saying nothing.

See also

Outage, Incident

Sources

Count this on a real site.

Watch my website