Uptime
Burn rate
Burn rate is how fast you are spending an error budget compared with the pace that would exactly use it up at period end. A rate of 1 empties the budget on the last day; a rate of 10 empties it in a tenth of the time.
How it is measured
Divide the observed bad-event rate in a window by the allowed rate. For a 99.9 percent SLO the allowed error rate is 0.1 percent. If the last hour shows 1.4 percent errors, burn rate is 14.
Alert on two windows: a short, fast one (5 minutes at about 14) to catch a hot incident, and a longer, slow one (6 hours at about 1) to catch a trickle. Pairing windows filters blips that clear on their own.
Worked example
A podcast host promises 99.9 percent over 30 days, a budget of 43.2 minutes. On day 6 a bad CDN rule makes 5 percent of requests fail for 40 minutes. 5 percent divided by 0.1 percent is a burn rate of 50. At that rate the whole month's budget would be gone in 14.4 hours.
The fast alert fires four minutes in. The rule is reverted at minute 40, having spent about 2 minutes of full-outage equivalent, or 4.6 percent of the month's budget. Left alone until morning, it would have burned all of it.
How it differs
Burn rate is the speed. Error budget is the tank. The budget says how much failure you may have this period, and burn rate says whether you are on course to run out. Budget excludes pace, and burn rate excludes how much is left in total.
Common errors
Using only one window. Forgetting to normalize by the SLO, so a 99.99 target and a 99.9 target share a threshold. Alerting on burn when traffic is tiny. Mixing calendar months with rolling 30 days. Not stating which events count as bad.
In practice
Set one fast and one slow burn alert on your main SLI, and let only the fast one page. Calibrate on last quarter's incidents by replaying them to see which would have fired. Silence low-volume routes where one error moves the ratio.