Uptime

Error budget

Error budget is the amount of failure an SLO allows in a period. A 99.9 percent monthly target leaves about 43 minutes.

How it is measured

Budget equals one minus the SLO, applied to the window. For request-based SLIs, 99.9 percent over 20 million requests leaves 20,000 bad requests. For time-based, over 30 days it leaves 43.2 minutes. Subtract what incidents have already consumed.

Show remaining budget as a number and a percentage, updated daily. A budget at 15 percent with 10 days left changes how risky a deploy is.

Worked example

A document-signing site holds 99.95 percent over 28 days, which is 20.2 minutes. By day 14 it has spent 6 minutes on a bad migration and 9 on a regional network fault. 15 of the 20.2 minutes are gone, and 5.2 remain for two more weeks.

The team defers a risky search rewrite to the next window and ships only fixes. Nothing broke because of the budget. The budget is why the release waited.

How it differs

Error budget is the allowance you can spend. SLO is the target that defines it. The SLO says 99.9, and the budget says 9 minutes left. The SLO excludes how much you have used so far. The budget excludes why that target was chosen.

Common errors

Treating it as a target to hit. Resetting it mid-month after a bad week. Counting events that users never saw. Never spending any, which usually means the SLO is too tight. Having no consequence when it runs out.

In practice

Pick a policy: below 25 percent remaining, freeze risky releases; at zero, only reliability work. Publish remaining budget where the whole team sees it. If you finish every month with 95 percent left, loosen the target or ship faster.

See also

SLO, Burn rate

Sources

Count this on a real site.

Watch my website