Uptime

Incident

Incident is a named event where service is hurt enough that people drop other work to fix it. It has an owner, a start, and an end.

How it is measured

Record severity, start, detection, mitigation, and resolution timestamps, and who was incident commander. Those timestamps give MTTD, MTTA, and MTTR for the event.

Declare by rule, not mood: for example SEV2 if checkout fails for more than 5 minutes or error budget burn exceeds 10 times. Without a threshold, long dull problems never get declared.

Worked example

At 19:48 on a Sunday, an image CDN purge job deletes the originals for a photography portfolio site. Thumbnails 404 on 1,200 pages. The on-call declares SEV2 at 19:55, opens a shared doc, and names Dana as commander and Luis as scribe. They restore from the bucket's versioning at 20:31.

Timeline: start 19:48, detected 19:52, declared 19:55, resolved 20:31, which is 43 minutes in all. The doc becomes the postmortem draft.

How it differs

Incident is the response unit with an owner and a clock. Outage is the user-facing condition of being down. You can have an incident with no outage, such as a security scare, and an outage so short that nobody declares an incident.

Common errors

No declared owner, so three people debug the same thing. Not writing a timeline while it happens. Declaring only when it is already bad. Fixing before communicating. Closing without a follow-up.

In practice

Write a one-page definition of severity levels and roles, then pin it where on-call looks. Run a pretend incident for 20 minutes. Make sure someone's job is to update the status page every 30 minutes.

See also

Outage, Postmortem

Sources

Count this on a real site.

Watch my website