A/B test calculator
Enter visitors and conversions for variants A and B. The A/B test calculator tells you whether the difference is statistically significant, shows the uplift and confidence interval, and works out how many visitors per variant a reliable test would need.
Example: Variant A: visitors 10,000 · Variant A: conversions 200 · Variant B: visitors 10,000 · Variant B: conversions 250 · Confidence level: 95% (standard) → Result: Significant at 95%: B converts better than A (p = 0.0171).. Source: NIST/SEMATECH e-Handbook of Statistical Methods, 7.3.3: comparing two proportions (pooled z-test). Updated: .
How it is calculated
How significance is calculated
The calculator uses the two-proportion z-test (NIST/SEMATECH e-Handbook): rates pA = conversions ÷ visitors, pooled proportion p = (cA + cB) ÷ (nA + nB) and z = (pB − pA) ÷ √(p(1 − p)(1/nA + 1/nB)). The two-sided p-value is the chance of seeing a difference at least this large if there were no real effect.
Example
| Visitors | Conversions | Rate | |
|---|---|---|---|
| A | 10,000 | 200 | 2.00% |
| B | 10,000 | 250 | 2.50% |
Uplift 25%, z = 2.38, p = 0.0171: significant at 95% confidence. The true difference lies between +0.09 and +0.91 percentage points with 95% confidence.
A/B test sample size
n per variant = (z1−α/2·√(2p̄(1 − p̄)) + z1−β·√(pA(1 − pA) + pB(1 − pB)))² ÷ (pB − pA)². Detecting 2.0% vs. 2.5% at 95% confidence and 80% power takes about 13,800 visitors per variant; 2.0% vs. 2.2% takes about 80,700.
Common mistakes
- Peeking: stopping as soon as results look significant inflates false positives.
- Underpowered tests: small uplifts need very large samples.
- Testing many variants or metrics without adjusting the significance level.
Frequently asked questions
How do you know if an A/B test is significant?
When the p-value is below your threshold, usually 0.05 (95% confidence). Then a difference this large is unlikely to be random.
How do you calculate A/B test significance?
With a two-proportion z-test: z = difference in conversion rates ÷ pooled standard error, then convert z into a p-value. The calculator does both.
How long should an A/B test run?
Until each variant reaches the sample size you calculated up front, and for at least one full week to smooth out weekday effects.
What sample size do I need for an A/B test?
It depends on your baseline rate and the uplift you want to detect. For 2.0% → 2.5% at 95% confidence and 80% power, about 13,800 visitors per variant.
What is statistical power?
The chance of detecting an effect that really exists. 80% is the usual target and is used here for the sample size.
Sources and legal basis
- NIST/SEMATECH e-Handbook of Statistical Methods, 7.3.3: comparing two proportions (pooled z-test)
- D. Zhang, ST 520 Statistical Principles of Clinical Trials, Ch. 6: Sample Size Calculations (two proportions)
As of:
Related tools
- Conversion rate calculatorCalculate your marketing conversion rate from visitors and conversions, plus average order value and revenue per visitor – or the traffic you need for a goal.
- P-Value CalculatorP-value calculator from a z-score, t-score or chi-square, or straight from a t-test with means, SDs and n – plus critical value and confidence interval.
- ROAS calculatorCalculate ROAS as a ratio and a percentage, plus ACOS, CPA and your break-even ROAS from your margin – using the Google Ads and Amazon Ads definitions.
- Customer lifetime value calculatorCalculate customer lifetime value (CLV/LTV) from order value, purchase frequency, lifespan or churn and margin – with CAC, LTV:CAC ratio and payback period.
- CPM, CPC and CTR calculatorCalculate CPM, CPC and click-through rate (CTR) from cost, impressions and clicks – or work backwards to budget, impressions and clicks from CPM and CTR.
- Break-even calculatorBreak-even point calculator for small businesses: units and sales dollars to break even from fixed costs, price and variable cost, plus a target profit.