Statistical Significance
Statistical significance is a judgment that an observed result, such as a difference between two versions in an A/B test, would be unlikely if there were no real effect. It is assessed by comparing a p-value against a threshold chosen in advance, called the significance level, which is commonly 0.05. A significant result suggests an effect is real. It does not show that the effect is large or important.
How Statistical Significance Works
Every significance test starts from a null hypothesis, the assumption that there is no real difference, for example that version A and version B convert at the same rate. The test then asks how surprising the observed data would be if that were true.
- The p-value is the probability of getting a result at least as extreme as the one observed, assuming the null hypothesis is true. A small p-value means the data would be unusual if there were no effect.
- The significance level (alpha) is the risk the team accepts of rejecting the null hypothesis when it is actually true, a false positive. An alpha of 0.05 means that when there is no real effect, the test will still wrongly flag one about 5% of the time. This matches the 95% confidence level often quoted in experiment tools.
- If the p-value falls below alpha, the result is called statistically significant and the null hypothesis is rejected.
- Statistical power is the probability of detecting an effect when one really exists. Power depends mainly on sample size and on how large the effect is, so small samples often miss real but modest effects.
Why Statistical Significance Matters
Product metrics fluctuate constantly. Without a significance test, teams cannot tell whether a lift in an A/B test is a real improvement or ordinary noise, and they end up shipping changes that do nothing or abandoning ones that work.
It is also widely misused. The American Statistical Association's 2016 statement on p-values sets out principles that apply directly to product work:
- A p-value is not the probability that the hypothesis is true, or the probability that the result happened by chance alone.
- Statistical significance does not measure the size of an effect or the importance of a result.
- Business decisions should not be based only on whether a p-value passes a threshold.
In practice, the most common error is checking results repeatedly and stopping as soon as they cross 0.05, or testing many metrics and reporting the one that did. Both make false positives far more likely than the stated 5%. The threshold, metric and sample size belong in the experiment design, decided before data arrives.
Statistical Significance Example
A team tests a new onboarding screen. After the planned sample is reached, 4.0% of control users and 4.4% of treatment users complete setup, and the test reports a p-value of 0.03. Because 0.03 is below the pre-set alpha of 0.05, the result is statistically significant: a difference this large would be unlikely if the screen had no effect. The team then asks the separate question of whether a 0.4 percentage point gain justifies the screen's maintenance cost, and looks at the confidence interval to see the plausible range of the true effect.
Statistical Significance vs. Practical Significance
Statistical significance asks whether an effect is likely to be real. Practical significance asks whether it is big enough to matter. With very large samples, even a tiny, commercially meaningless difference can be statistically significant. With small samples, a large and valuable effect may fail to reach significance. Good decisions consider both, alongside the effect size and the original hypothesis.