Uptime
Outage
Outage is a period when a service is down or unusable for the people who need it. Slowness and partial failure get softer words.
How it is measured
Mark it when a defined critical function fails for real users: homepage 5xx, checkout error, login refused. Use multi-region probes and server logs. A single failed probe is a suspect. Three consecutive failures from two regions is a fact.
Record scope too: all users, one country, one plan. An outage for 8 percent of users is still an outage, and the number belongs in the record.
Worked example
A Danish language-learning site loses its only load balancer to a provider reboot at 07:14 local time. Every region fails, with 100 percent of probes red. The provider restores the box at 07:52.
That is 38 minutes of outage during peak morning study time, and 1,900 users saw a connection error. In the status page history it is a clean red bar, not a debate.
How it differs
Outage is a hard failure. Incident is the managed response, which can cover an outage, a brownout, or a near miss. Outage describes the user's experience. Incident describes the team's work.
Common errors
Calling slowness an outage. Calling it maintenance after the fact. Counting only full-site failures. Reporting from a single probe location. Not recording partial ones. Reporting duration from the alert instead of the start.
In practice
Agree which failing function counts as an outage for your site, and wire a probe to it. Publish times and scope after each one.