Sample size in A/B testing and five common misconceptions
The sample needed per group from power, significance level and minimum detectable difference, the 5% to 6% example, and misconceptions such as stopping early and multiple comparisons; a short note on survey samples.
An A/B test compares two versions using randomly split users. Tests often fail for the wrong reason: they are run for much shorter than needed, or started with too small a sample to see a small difference. Calculating the sample size at the start largely prevents both problems.
What does the sample size depend on?
Four quantities determine one another:
- **The current rate** (for example a 5% conversion rate).
- **The smallest difference you want to catch** (for example from 5% to 6%).
- **The significance level** (usually 5%, two-sided): the risk of wrongly saying "there is a difference" when there is none.
- **The power** (usually 80%): the probability of catching a difference when it really exists.
Tightening any one of them enlarges the sample. With the two-proportion z-test approach, examples:
- 5% → 6% (a 20% relative lift): about 8,158 users per group, about 16,300 in total for two groups
- 5% → 7%: about 2,213 per group
- 5% → 5.5% (a 10% relative lift): about 31,234 per group
When the difference is halved, the sample grows about fourfold: the sample is inversely proportional to the square of the difference. That is why measuring small improvements is expensive.
Five common misconceptions
- **Stopping the test when it turns significant.** Looking at the result continuously and stopping at the first significant p-value raises the false-positive rate. Write down the sample size and duration beforehand and do not decide before reaching them (or use sequential testing methods).
- **Trying many metrics or segments.** Out of twenty metrics, one is expected to come out significant by chance. Fix the primary metric in advance and treat the others as exploration.
- **Skipping the full weekly cycle.** Weekday and weekend behaviour differ. Run the test for at least one, preferably two full weeks.
- **Confusing significance with effect size.** In a very large sample even a trivial difference turns significant; look at the confidence interval and the business value.
- **Not noticing a broken split.** If the share of users between groups deviates from plan (for example 52-48 instead of 50-50), there may be an assignment error; the result should not be interpreted without this check.
Survey samples: a short note
In surveys the same logic works with different quantities: margin of error, confidence level and population. For 95% confidence and a ±5% margin of error, in the worst case (proportion 50%) the sample needed for an infinite population is 384.15, which rounds up to 385. If the population is 2,000 people, the finite population correction brings it down to 323. With a 30% response rate you must invite 323 / 0.30 ≈ 1,077 people. Whether respondents represent the whole population (response bias) is a separate problem; sample size does not solve it.
Try it with the tools
You can calculate the needed sample and the significance of a result in the A/B test sample size and significance tool: the two-proportion z-test, p-value, confidence interval and relative lift come out in your browser. For surveys there is the survey sample size tool, and for differences in means the t-test calculator. To review your experiment design together, get in touch.
Let's discuss this for your plant
More notes
All notes →- 3 min readSix common mistakes on medical device UDI labelsPackaging levels, changes that require a new UDI-DI, confusing Basic UDI-DI with UDI-DI, missing production identifiers, labels that don't match the barcode and unmeasured print quality: the most common UDI mistakes for medical devices.
- 3 min readSix ways the case–pallet relationship breaks in aggregationRejected packs, partial cases, sampling, manual handling, double assignment and pallet breakdown: the most common situations in which aggregation records drift from physical content, and how to prevent them on the line.
- 2 min readGaps and bad records in SCADA data: what to do before modellingTimestamps, gaps, frozen sensors, physically impossible values, curtailment and maintenance periods, sensor replacement: what to check in SCADA data before building a predictive maintenance or power forecasting model.