Analysis

Statistical significance

Also called p-value threshold.

Statistical significance is a test result showing that a difference is unlikely under a stated null at a chosen alpha. It says the gap is probably not noise, not that it is big or important.

How it is measured

You pick alpha, often 0.05, before the run. The test returns a p-value, the probability of seeing data at least this extreme if there were no true difference. If p is below alpha, you reject the null.

Report the effect size and an interval as well as the p-value. A tiny difference can be significant with enough traffic and still not matter.

Worked example

A ferry-booking site tests a new fare table. Over 41,000 visits, the new layout converts at 5.30 percent against 5.15 percent. The p-value is 0.04.

The test passes but the lift is 0.15 points, about 60 extra bookings a month. The engineering cost to roll it out is more than that is worth, so the site leaves the old layout alone.

How it differs

Statistical significance says a difference is probably not chance. A confidence interval shows how large it might be. The first gives a verdict; the second gives a range of sizes.

Common errors

Reading p of 0.04 as 96 percent sure. Peeking at the result and stopping when p dips. Testing twenty things and celebrating the one that passed. Equating significant with important. Dropping alpha after the fact.

In practice

Choose alpha and run length before the test. Report the size of the effect, not only the verdict. Treat a result that barely passes with care and repeat it if the decision is expensive.

See also

Confidence interval, Sample size, A/B test

Sources

Count this on a real site.

Watch my website