A/B Testing
A/B testing, also called split testing, is an online controlled experiment that randomly assigns users to two versions of a product experience: the current version (the control, A) and a changed version (the treatment, B). The team compares a predefined metric between the two groups, and statistical analysis shows whether the difference is likely caused by the change rather than by chance.
How A/B Testing Works
An A/B test has a few essential parts:
- Variants. The control and one or more treatments. A test with several treatments is often called an A/B/n test.
- Randomization unit. The entity that is randomly assigned, usually a user. Assignment should be persistent, so the same user sees the same variant on every visit.
- Primary metric. The measure that decides the test, such as purchases per user. Ron Kohavi and colleagues call this the Overall Evaluation Criterion (OEC).
- Guardrail metrics. Measures that must not get worse, such as page load time, error rate or refunds.
- Sample size and duration. Calculated before launch so the test has enough statistical power to detect the smallest effect worth acting on.
Because users are assigned at random, the two groups are alike on average in everything except the change. Any reliable difference in the metric can therefore be attributed to the change. The result is judged with a statistical test, and a difference that passes the pre-set threshold is called statistically significant.
Why A/B Testing Matters
A/B testing is the most direct way to measure the causal effect of a product change on real users. That matters because intuition is a weak guide: in a 2012 paper, Kohavi and colleagues reported that only about one third of ideas tested at Microsoft improved the metrics they were designed to improve.
Results can mislead when the test is run carelessly. Common problems include stopping the test as soon as the numbers look good (called peeking), checking dozens of metrics until one appears significant, and a sample ratio mismatch, where the split between groups differs from the plan, often a sign of a bug. Mature teams also run A/A tests, where both groups get the same experience, to check that the system does not report false differences.
A/B testing also has limits. It needs enough traffic to reach a reliable answer, it shows which version wins rather than why, and it suits optimizing an existing product better than deciding what new product to build.
A/B Testing Example
An online store wants to know whether a single-page checkout beats its current three-step checkout. Each visitor who reaches checkout is randomly assigned to one version and stays in it. The primary metric is completed orders per user. Guardrails are average order value, payment errors and page load time. The team calculates the sample it needs, runs the test for two full weeks to cover weekday and weekend behavior, and does not act on early results. At the end, the single-page version shows a statistically significant increase in completed orders with no guardrail harmed, so it is rolled out to everyone.
The change in that example was framed in advance as a hypothesis with an expected direction, which is what keeps the analysis honest.