Uptime

SLI

Also called service level indicator.

SLI is a measured number that describes how a service is performing from a user's view, such as the fraction of requests that succeed in under 500 ms.

How it is measured

Pick a good-events over valid-events ratio: successful requests over all valid requests, probes under threshold over all probes, or syncs finished under 5 minutes over all syncs. Ratios between 0 and 100 percent work with SLOs.

Choose where to measure: at the load balancer, a client, or a synthetic probe. Each sees different failures. Say which one and over what window.

Worked example

A bookshop's search API logs 2.4 million requests in a week. 2,397,600 return non-5xx, and 2,381,000 of the total are under 800 ms. The availability SLI is 99.90 percent. The latency SLI is 99.21 percent.

The two SLIs tell different stories: the site rarely fails, but 0.8 percent of searches are slow. Each gets its own SLO instead of one blended score.

How it differs

SLI is the measurement. SLO is the target for it. The SLI excludes the goal. The SLO excludes the raw number.

Common errors

Measuring server CPU instead of user outcomes. Including bots and health checks. Using averages. Choosing too many. Mixing windows. Measuring at a layer that misses failures, such as the app instead of the CDN.

In practice

Choose two SLIs: one for availability and one for latency on your most important path. Define good and valid events in one sentence each. Compute them for the last 30 days before you choose a target.

See also

SLO, Availability

Sources

Count this on a real site.

Watch my website