A/B Test Statistical Significance Calculator
You changed a headline, a form, or an ad — and the numbers went up. One question matters: is that a real gain or random noise? Enter both variants below and the calculator answers it. Every term is explained in plain English underneath.
Enter your test data
Visitors — how many people saw the variant. Conversions — how many of them did the thing you wanted (form, call, purchase).
Enter your data
How to use it in 3 steps
- Pull visitors and conversions for each variant from your analytics, over the exact same period.
- Type in the four numbers. It calculates as you type — nothing to click.
- Read the confidence level. 95% or higher means the difference is proven. Below that, you don't have enough data yet.
How to read the result
- 95% and above
- The difference is almost certainly real. You can ship variant B — after checking that the gain is worth the work.
- 90% to 95%
- It looks real but isn't proven. Keep the test running; do not call it now.
- Below 90%
- Not enough data. This does NOT mean there is no difference — it means the test hasn't answered yet.
Every term in plain English
- Conversion rate
- The share of people who did what you wanted. 40 leads from 1,000 visitors is 4%.
- Variant A and variant B
- A (the control) is what already runs. B is what you're testing. The comparison only holds if both ran over the same period on comparable traffic.
- Uplift
- How much better or worse B is than A, in relative percent. Going from 4% to 5% is a 25% uplift, not a 1% uplift.
- Percentage point (pp)
- The gap between 4% and 5% is 1 percentage point — and also a 25% relative uplift. The same number gets reported whichever way flatters the report, so always check which one you're being shown.
- Statistical significance
- The answer to one question: could this gap have happened by pure chance? If it almost certainly couldn't, the difference is significant.
- p-value
- The probability of seeing a gap this large (or larger) if there were genuinely no difference. p = 0.03 means a 3% chance that randomness fooled you.
- Confidence level
- 100% minus the p-value. The standard bar is 95%. Below it, a result isn't treated as proven.
- Confidence interval
- The range the true difference sits in. If it reads "−0.2 to +2.4 pp", variant B might actually be worse — zero is inside the range.
- False positive
- The test says there's a difference when there isn't. At a 95% bar this happens roughly 1 time in 20 — so one winning test out of twenty is normal, not a discovery.
- Sample size
- How many people the test needs. The smaller the difference you want to detect, the more traffic it takes: a 50% lift shows up in hundreds of visits, a 5% lift needs tens of thousands.
- Peeking
- The habit of stopping a test the moment the number first turns green. Do that and you can "win" almost any test — including a variant tested against itself. Decide the duration and volume up front.
- Practical significance
- Even a mathematically proven 0.3% gain may not pay for the work or the risk. Significance is about math, not money. The money math is a separate calculation.
- Two-tailed test
- This calculator checks the difference in both directions: B can come out better or worse than A. That's more honest than a one-tailed test, which assumes up front that the new version wins.
Common mistakes that ruin a test
- Calling the test at 200 visitors because "you can already see it".
- Counting clicks instead of leads and revenue. A more clickable variant can easily earn less.
- Changing two things at once — then you can't tell which one worked.
- Running for less than a week: weekdays and weekends behave differently.
- Reading "not significant" as "no difference". It only means "not enough data".
What the calculator computes
A two-tailed z-test for two proportions: the test statistic uses the pooled conversion rate, the p-value comes from the normal distribution, and the 95% confidence interval covers the absolute difference in conversion rates. This is the standard method for A/B tests with a binary outcome (converted or not). With small samples — fewer than 5 conversions in either variant — the calculator warns you that the math isn't trustworthy yet.
Get an Express Growth Audit for Your Business
Enter your website URL — we'll send a breakdown of marketing leaks within 24 hours.