Uptime

Downtime

Downtime is the stretch of time when a service did not answer or did not work as agreed. It is counted in minutes, and it starts when users were hurt, not when someone noticed.

How it is measured

Mark the start at the first failed probe in a confirmed run, or the first user failure in logs, and the end at the first sustained success. Subtract from the period to get availability. Keep the raw timestamps, not just the duration.

Decide up front what counts: full failure only, or brownouts above a threshold too; scheduled maintenance in or out. With 1-minute probes and a 3-failure confirmation rule, reported downtime can lag the real start by two minutes.

Worked example

A small SaaS invoicing app has a TLS certificate expire at 00:00 UTC on 3 March. Browsers show a warning and API clients fail the handshake. Detection waits for the 3-failure rule, and someone acknowledges at 00:07. The certificate is replaced by 00:41.

The incident record says downtime ran 00:00 to 00:41, which is 41 minutes, not 34, because the failure began at midnight. That single event is 95 percent of a 43.2-minute monthly allowance.

How it differs

Downtime is the bad stretch itself. Uptime is the proportion of time that was not bad. Downtime is an absolute length you can attach to one incident. Uptime is a ratio over a period and hides how many separate pieces made it.

Common errors

Starting the clock when the alert arrived. Excluding partial outages without a rule. Rounding every incident down to five minutes. Counting only what one region's monitor saw. Not recording the end.

In practice

For each incident this month, write start and end from probe logs, not memory. Keep one agreed definition of what counts and apply it every time. Add up minutes per quarter and compare them with what you promised.

See also

Uptime, Incident

Sources

Count this on a real site.

Watch my website