Analysis

Anomaly detection

Also called spike detection.

Anomaly detection is a method that flags values outside a band built from past behavior. The band is the product; a flag is only as good as the history behind it.

How it is measured

The common approach compares a value against a rolling mean and spread for the same hour or weekday, often a fixed number of standard deviations. You need enough history to cover at least a few full cycles of the weekly shape.

Check the detector by replaying the last quarter. List every flag it would have raised and mark which ones matched a real event in your notes. Count the misses too, since a detector that never fires has no false alarms and no use.

Worked example

A language-learning site averages 1,420 signups on weekdays and 610 on weekends. A detector built on a flat all-week mean flags every Saturday as a collapse, six weekends in a row.

Rebuilding the band per weekday gives a weekend range of 540 to 690 and the false flags stop. A real drop to 210 on a Wednesday in June is then caught within the hour.

How it differs

Anomaly detection judges whether a value is odd. An alert is the delivery step that tells a person. You can send alerts from fixed thresholds without any detection model.

Common errors

Training the band on a period that included an outage. Ignoring weekday and holiday shape. Treating every flag as a fault. Using a window too short to learn a pattern. Never tuning after false alarms.

In practice

Start with one metric you care about and a band per weekday. Review flags weekly for a month and mark each real or false. Only wire the flag to a pager once most of them are real.

See also

Alert, Seasonality, Moving average

Sources

Count this on a real site.

Watch my website