Tool
A/B Test & Lift Calculator
Significance and revenue impact in one tool. Two-proportion z-test, p-values, conversion rate lift, and projected revenue impact — paste your numbers, see if your test is real.
Your Test Data
Used to calculate projected revenue impact
Confidence level
Use 99% for high-stakes changes (pricing, checkout)
Enter visitors and conversions for both variants to see significance
Statistical Significance
How to read your result
Three patterns — three different next steps.
Not significant?
Don't ship. Underpowered tests look like wins randomly — you're seeing noise, not a real effect. Either run longer to gather more data, increase the minimum detectable effect, or accept the null and move on.
Marginally significant?
Be careful. p=0.05 means roughly 1 in 20 tests is a false positive. For high-stakes changes — pricing, checkout, anything you can't reverse — hold to 99% confidence before you ship. The bar should match the cost of being wrong.
Strongly significant?
Ship it — but track the metric for 4+ weeks post-launch. Some lifts decay (novelty effect), and what looks like a winner in week 1 can flatten out by week 6. Confirm it holds before you trust it as a baseline.
Frequently asked questions
What does p-value actually mean?
P-value is the probability of seeing a lift this large by random chance if the variants were truly identical. p=0.05 means there's a 5% chance the result is noise, not a real effect. The lower the p-value, the stronger the evidence that your variant truly outperforms the control.
95% vs 99% confidence — when should I use which?
Use 95% confidence for routine UI and copy tests where the cost of being wrong is low. Use 99% for changes that affect pricing, checkout flow, or anything you can't reverse easily — the higher bar reduces the chance of shipping a false positive on a high-stakes change.
Why does my test "go non-significant" the longer I run it?
Reverting confidence usually means the early lift was noise, not signal. Small samples produce volatile conversion rates. Trust the longer-running result — and avoid stopping a test the moment it crosses significance, which inflates false-positive rates substantially.
What sample size do I need?
Sample size depends on baseline conversion rate and minimum detectable effect (MDE). Tools like Optimizely's calculator give you a pre-test number. This calculator shows post-hoc significance — what your current data tells you after the test has run.
How does Pace use A/B test data?
Pace doesn't run on-site A/B tests — but it tracks ad-creative performance and surfaces fatigued creatives (Meta: frequency above 3 with CTR below 1%) so you know exactly when to rotate. The same statistical thinking that powers this calculator informs how Pace's Creative Lens flags ads.
Want creative rotation tied to performance — not your gut feel?
Pace tracks creative-level CTR, CPA, and ROAS across Meta, TikTok, and Google — and surfaces creatives showing fatigue or runaway lift before the data goes stale.
| This calculator | With Pace | |
|---|---|---|
| Frequency | One-shot, when you remember | Daily significance checks |
| Channels | One test at a time | Meta, TikTok, and Google creatives |
| Action | Tells you the p-value | Surfaces winners and laggards |
| History | Resets on refresh | Rolling creative-fatigue trail |