Uptime

Rollback

Rollback is undoing a release by returning to the last known good version. You do it when the new one hurts users or the SLO.

How it is measured

Measure time from decision to old version serving, and from regression start to decision. Include database state: can the schema go back? Track rollback rate per release as a quality signal.

Set triggers before shipping: error rate above 2 percent for 5 minutes, a checkout drop, or burn rate above 10. Pre-set numbers make the decision a rule instead of a debate.

Worked example

A theme update on a WordPress store goes live at 16:00 and breaks the mini-cart on mobile. Add-to-cart events fall from 180 per hour to 40 by 16:30. The owner decides at 16:34, restores the previous theme folder, and flushes the cache by 16:41.

The cost was 41 minutes of weak mobile sales. A trigger of add-to-cart under half of the same hour over the last four weeks would have flagged it at 16:12.

How it differs

Rollback returns to a previous version after harm. Canary limits who sees a new version before harm spreads. Rollback excludes prevention. Canary excludes recovery once everyone has the build.

Common errors

Releasing migrations that cannot revert. Deleting old artifacts. Rolling back without telling anyone. Debating for 30 minutes. Not rolling back config along with code. Rolling forward when a rollback is faster.

In practice

For your next deploy, write the rollback command and test it on staging. Keep the last three builds. Agree on a number that triggers a rollback.

See also

Canary, Blue-green deploy

Sources

Count this on a real site.

Watch my website