Skip to content

A/B test sample size and significance calculator

See how many visitors you need before you start, and whether the difference is more than chance once the test ends. There are two modes: required sample size and result significance. Formulas and common mistakes are below.

Free tool · Data analytics
What to calculate?
Type of effect
Test direction

The calculation runs in your browser; the values you enter are not sent anywhere.

Required sample per group
31,234
Total (both groups)
62,468
Expected rate in group B
5.5%
Absolute difference
+0.5 points
Relative difference
+10%
Estimated duration
32 days
z (significance)
1.96
z (power)
0.842

If the minimum effect changes

If the minimum effect changes
EffectB ratePer group
+5%5.25%122,124
+7.5%5.38%54,903
+10%5.5%31,234
+15%5.75%14,193
+20%6%8,158

Halving the effect roughly quadruples the sample you need.

n = ( z₁₋α/₂·√(2·p̄·q̄) + z₁₋β·√(p₁q₁ + p₂q₂) )² / (p₁ − p₂)², p̄ = (p₁ + p₂)/2 (per group)

The results rely on the normal approximation for two proportions and assume the test is completed with the sample size fixed in advance. They are not sufficient on their own for a decision; experiment design, multiple comparisons and business impact need separate assessment.

Shall we review your experiment design, measurement plan and how results feed into decisions together?

Request a call

01

How to use it

  1. A

    Before the test, pick the "Required sample size" mode: enter the baseline rate, the smallest effect you want to see, α and power; read the visitors needed per group and the estimated duration.

  2. B

    Run the test until it reaches that sample; do not look at interim results and stop early.

  3. C

    When it ends, switch to "Result significance" and enter visitors and conversions for both groups; read the p-value and the confidence intervals of the difference and relative lift.

02

What does the sample size depend on?

Four things determine the sample you need: the baseline rate, the smallest effect you want to detect (the minimum detectable effect), the significance level α and the statistical power. As the effect shrinks the sample grows very fast: halving the effect roughly quadruples the visitors needed, because the required sample is inversely proportional to the square of the effect. For the same relative effect, the lower the baseline rate, the larger the sample.

The tool computes the per-group sample with the usual normal approximation for a two-proportion test: n = (z₁₋α/₂·√(2·p̄·q̄) + z₁₋β·√(p₁q₁ + p₂q₂))² / (p₁ − p₂)². Here p₁ is the baseline rate, p₂ the expected new rate, p̄ their mean and q = 1 − p. Groups are assumed equal in size and the result is rounded up. This is the formula used by R's power.prop.test.

The minimum effect can be entered as relative or absolute. A 10% relative effect on a 5% baseline means moving to 5.5% (0.5 percentage points); in absolute mode you would enter percentage points directly. Mixing the two changes the sample many times over, so the calculator always shows the expected B rate separately.

03

What do the p-value and confidence interval say?

The p-value answers: if the two versions were truly identical, how likely is a difference this large or larger? If p < α the difference counts as statistically significant. The p-value does not say the difference is large or matters to the business; with a very large sample a trivial difference can be significant. A non-significant result does not mean "no difference" either; the data may be too thin to show one.

The confidence interval tells you more: it shows the range of plausible differences. If the interval contains zero, the difference cannot be told apart from zero at that confidence. The interval for the difference uses the unpooled standard error; for the relative lift the Katz log-ratio method is used. The confidence level is 1 − α; "there is a 95% chance the true difference lies in this interval" is not quite the right reading, the correct one being that if the same method were repeated many times, about 95% of the intervals would contain the true difference.

The test is a z-test: the group rates are pooled to obtain the standard error. If the expected number of converters or non-converters in either group is below 5, the normal approximation weakens and the tool shows a warning.

04

Common mistakes in A/B tests

Peeking and stopping early: looking at the p-value before the sample is reached and ending the test at the first significant reading raises the false-positive rate considerably. Fix the sample in advance or use sequential testing methods. Trying many variants or metrics at once causes the same problem; each comparison gives an extra chance, so a correction (for example Bonferroni) is needed.

Sample composition: if the mechanism that splits traffic is broken (sample ratio mismatch) the two groups cannot be compared. Check that group sizes match the intended ratio. Also run the experiment for at least one full business cycle (weekly pattern); a novelty effect can inflate results in the first days.

Confusing significance with importance: check whether the lower end of the confidence interval still shows an improvement that matters to the business. With a small sample the interval is wide; read it together with the minimum effect and the cost before deciding.

FAQ

Are the numbers I enter sent anywhere?
No. All calculations run in code inside your browser; no value is transmitted to a server or stored.
What should I choose as the minimum detectable effect?
Choose the smallest improvement that would create value for the business, for example a 5% relative increase in conversion. If collecting the visitors needed to detect a smaller effect is costly, design the test for a larger effect or allow more time.
Why are power and α set to 80% and 5%?
They are established defaults, not a law of nature. If a false positive is costly, lower α (for example 1%); if missing an effect is costly, raise the power (for example 90%). Both increase the sample you need.
When is a one-sided test appropriate?
Only if you care solely about "is B better than A?" and B being worse would lead to the same decision (for example not shipping it). Decide before the test; switching to one-sided after seeing the result makes the p-value unfairly small. If unsure, use two-sided.
How is the test duration worked out?
We divide the total sample by the combined daily visitors of both groups and round up. The real duration depends on traffic swings, the weekly pattern and the share of users eligible for the test; plan at least one full week or business cycle.
The result is significant; should I ship right away?
The p-value alone is not enough. Also look at the lower end of the confidence interval, your measurement plan (did you try several metrics or variants), sample ratio mismatch and the business impact.

Let's make your experiment results trustworthy

We work through your data with you on experiment design, the measurement plan and linking results to decisions. Let's discuss your needs in a free discovery call.