A/B test calculator

Enter visitors and conversions for variants A and B. The A/B test calculator tells you whether the difference is statistically significant, shows the uplift and confidence interval, and works out how many visitors per variant a reliable test would need.

Example: Variant A: visitors 10,000 · Variant A: conversions 200 · Variant B: visitors 10,000 · Variant B: conversions 250 · Confidence level: 95% (standard) → Result: Significant at 95%: B converts better than A (p = 0.0171).. Source: NIST/SEMATECH e-Handbook of Statistical Methods, 7.3.3: comparing two proportions (pooled z-test). Updated: .

My ToolboxYour inputs are saved in this browser only.

Result

Result
Significant at 95%: B converts better than A (p = 0.0171).
Conversion rate A
2%
Conversion rate B
2.5%
Uplift of B over A
25%
p-value (two-sided)
p = 0.0171 · z = 2.38
Confidence interval of the difference
+0.09 to +0.91 percentage points (95%)
Visitors needed per variant (80% power)
13,809
Statistical estimate only. Don’t stop a test the moment p drops below 5% – fix the sample size in advance.

How it is calculated

How significance is calculated

The calculator uses the two-proportion z-test (NIST/SEMATECH e-Handbook): rates pA = conversions ÷ visitors, pooled proportion p = (cA + cB) ÷ (nA + nB) and z = (pB − pA) ÷ √(p(1 − p)(1/nA + 1/nB)). The two-sided p-value is the chance of seeing a difference at least this large if there were no real effect.

Example

VisitorsConversionsRate
A10,0002002.00%
B10,0002502.50%

Uplift 25%, z = 2.38, p = 0.0171: significant at 95% confidence. The true difference lies between +0.09 and +0.91 percentage points with 95% confidence.

A/B test sample size

n per variant = (z1−α/2·√(2p̄(1 − p̄)) + z1−β·√(pA(1 − pA) + pB(1 − pB)))² ÷ (pB − pA)². Detecting 2.0% vs. 2.5% at 95% confidence and 80% power takes about 13,800 visitors per variant; 2.0% vs. 2.2% takes about 80,700.

Common mistakes

Frequently asked questions

How do you know if an A/B test is significant?

When the p-value is below your threshold, usually 0.05 (95% confidence). Then a difference this large is unlikely to be random.

How do you calculate A/B test significance?

With a two-proportion z-test: z = difference in conversion rates ÷ pooled standard error, then convert z into a p-value. The calculator does both.

How long should an A/B test run?

Until each variant reaches the sample size you calculated up front, and for at least one full week to smooth out weekday effects.

What sample size do I need for an A/B test?

It depends on your baseline rate and the uplift you want to detect. For 2.0% → 2.5% at 95% confidence and 80% power, about 13,800 visitors per variant.

What is statistical power?

The chance of detecting an effect that really exists. 80% is the usual target and is used here for the sample size.

Sources and legal basis

As of:

Related tools