Uptime

Degraded

Degraded is a state where a service works in part: some features, regions, or users fail while others are fine. It is the label for not down, not healthy.

How it is measured

Mark degraded when a named component misses its target but the whole does not: search returning errors, one region above 3 seconds, image uploads failing. Component-level probes make it visible. A single homepage check never does.

Express it as a share: 2 of 5 regions failing, 1 of 6 endpoints red, 8 percent of requests erroring. Status pages usually show degraded as a middle color between operational and outage.

Worked example

A job board's API serves search, apply, and employer login. At 15:30 the search index node dies. /search returns 503 while /apply and /login keep working. Probes show 2 of 3 critical paths green.

The team posts Degraded: search on the status page, not Outage. Applications still arrive, 340 in the hour. Once search returns at 16:05, the incident record lists 35 minutes of degradation that counted against the search SLI only.

How it differs

Degraded names a part that is wrong and leaves the working parts out of it. Brownout describes general slowness or partial failure across the whole path under load. Degraded excludes the pieces that still work. Brownout excludes any clean boundary around what broke. Neither is a full outage.

Common errors

Using it as a catch-all so nobody reads it. Having no per-component probes. Not saying which feature is affected. Counting it as full downtime or as zero. Leaving the label up after recovery.

In practice

Split your probes by feature and region, and agree what share of failure moves a component to degraded. Show it on the status page with the affected part named. Decide in your SLO whether degraded minutes spend error budget, and by how much.

See also

Brownout, Error budget

Sources

Count this on a real site.

Watch my website