Skip to content

Sample size in A/B testing and five common misconceptions

The sample needed per group from power, significance level and minimum detectable difference, the 5% to 6% example, and misconceptions such as stopping early and multiple comparisons; a short note on survey samples.

3 min readA/B testing · Sample size · Statistics · Survey

An A/B test compares two versions using randomly split users. Tests often fail for the wrong reason: they are run for much shorter than needed, or started with too small a sample to see a small difference. Calculating the sample size at the start largely prevents both problems.

What does the sample size depend on?

Four quantities determine one another:

  • **The current rate** (for example a 5% conversion rate).
  • **The smallest difference you want to catch** (for example from 5% to 6%).
  • **The significance level** (usually 5%, two-sided): the risk of wrongly saying "there is a difference" when there is none.
  • **The power** (usually 80%): the probability of catching a difference when it really exists.

Tightening any one of them enlarges the sample. With the two-proportion z-test approach, examples:

  • 5% → 6% (a 20% relative lift): about 8,158 users per group, about 16,300 in total for two groups
  • 5% → 7%: about 2,213 per group
  • 5% → 5.5% (a 10% relative lift): about 31,234 per group

When the difference is halved, the sample grows about fourfold: the sample is inversely proportional to the square of the difference. That is why measuring small improvements is expensive.

Five common misconceptions

  • **Stopping the test when it turns significant.** Looking at the result continuously and stopping at the first significant p-value raises the false-positive rate. Write down the sample size and duration beforehand and do not decide before reaching them (or use sequential testing methods).
  • **Trying many metrics or segments.** Out of twenty metrics, one is expected to come out significant by chance. Fix the primary metric in advance and treat the others as exploration.
  • **Skipping the full weekly cycle.** Weekday and weekend behaviour differ. Run the test for at least one, preferably two full weeks.
  • **Confusing significance with effect size.** In a very large sample even a trivial difference turns significant; look at the confidence interval and the business value.
  • **Not noticing a broken split.** If the share of users between groups deviates from plan (for example 52-48 instead of 50-50), there may be an assignment error; the result should not be interpreted without this check.

Survey samples: a short note

In surveys the same logic works with different quantities: margin of error, confidence level and population. For 95% confidence and a ±5% margin of error, in the worst case (proportion 50%) the sample needed for an infinite population is 384.15, which rounds up to 385. If the population is 2,000 people, the finite population correction brings it down to 323. With a 30% response rate you must invite 323 / 0.30 ≈ 1,077 people. Whether respondents represent the whole population (response bias) is a separate problem; sample size does not solve it.

Try it with the tools

You can calculate the needed sample and the significance of a result in the A/B test sample size and significance tool: the two-proportion z-test, p-value, confidence interval and relative lift come out in your browser. For surveys there is the survey sample size tool, and for differences in means the t-test calculator. To review your experiment design together, get in touch.