Uptime

Status page

Status page is a public page that says which parts of your service are working and what you are doing about problems.

How it is measured

It is not a metric, so check its content: component list, current state (operational, degraded, outage), incident history, timestamps, and a subscribe option. Measure update lag from incident declared to first post, with a target of 10 minutes.

Host it away from your infrastructure. A status page on the same server as the site goes down with it.

Worked example

An e-commerce plugin vendor has components for API, Dashboard, Webhooks, and Email. At 09:20 webhook deliveries start failing. The status page flips Webhooks to Degraded at 09:26 with the line Some deliveries delayed; retries are queued. Updates follow at 09:50 and 10:15, and it resolves at 10:40.

Support tickets that morning: 12, rather than the usual 60 for events like this. The page also shows the past 90 days so a customer can verify a claim.

How it differs

Status page tells people what is going on now. Incident is the internal event with a commander and a timeline. The status page excludes the debugging detail. The incident record excludes the public tone.

Common errors

Hosting it on the same infrastructure. Updating late. Vague text like some issues. Never marking things degraded. Deleting history. Relying on manual updates only, so it lags.

In practice

Create the page on a separate provider, list five components, and connect your probes to flip states. Agree who posts the first update. Write three template messages in advance.

See also

Incident, Scheduled maintenance

Sources

Count this on a real site.

Watch my website