Uptime
SLO
Also called service level objective.
SLO is the internal target for an SLI over a period, such as 99.9 percent of requests succeeding over 30 days. It is the number a team steers by.
How it is measured
State the SLI, target, and window: 99.9 percent of valid requests return non-5xx in under 1 second, measured over a rolling 28 days at the load balancer. The target defines the error budget.
Pick from history and users' tolerance, not aspiration. Start slightly below what you achieved last quarter, and review quarterly.
Worked example
A fitness-class booking app reviews last quarter's data: availability by month was 99.97, 99.81, and 99.93 percent. The team sets the SLO at 99.9 over 28 days, which gives 40.3 minutes of budget. In the first window, a 22-minute outage plus several short slow patches use 31 minutes.
With 9 minutes left, the team delays a schema change by a week. A year earlier they would have shipped it anyway.
How it differs
SLO is the internal target. SLA is the external contract with penalties. An SLO excludes legal remedy and can be missed with a postmortem. An SLA excludes the internal margin.
Common errors
Targeting 100 percent. Keeping many SLOs no one watches. Choosing the target before measuring. Having no consequence on a miss. Using calendar-month windows that reset the pain. Setting it equal to the SLA.
In practice
Write one SLO for your key user path with SLI, target, and window. Compare it to six months of data. Decide what you do when budget falls below 25 percent.