Skip to content

Tool

A/B Test & Lift Calculator

Significance and revenue impact in one tool. Two-proportion z-test, p-values, conversion rate lift, and projected revenue impact — paste your numbers, see if your test is real.

Your Test Data

$

Used to calculate projected revenue impact

Confidence level

Use 99% for high-stakes changes (pricing, checkout)

Enter visitors and conversions for both variants to see significance

Statistical Significance

-- conf · p=--
Enter your numbers
CR Variant A 0%
CR Variant B 0%
Absolute Lift 0pp
Relative Lift % 0%
P-value --
Samples Needed / Variant --

How to read your result

Three patterns — three different next steps.

Not significant?

Don't ship. Underpowered tests look like wins randomly — you're seeing noise, not a real effect. Either run longer to gather more data, increase the minimum detectable effect, or accept the null and move on.

Marginally significant?

Be careful. p=0.05 means roughly 1 in 20 tests is a false positive. For high-stakes changes — pricing, checkout, anything you can't reverse — hold to 99% confidence before you ship. The bar should match the cost of being wrong.

Strongly significant?

Ship it — but track the metric for 4+ weeks post-launch. Some lifts decay (novelty effect), and what looks like a winner in week 1 can flatten out by week 6. Confirm it holds before you trust it as a baseline.

Frequently asked questions

What does p-value actually mean?

P-value is the probability of seeing a lift this large by random chance if the variants were truly identical. p=0.05 means there's a 5% chance the result is noise, not a real effect. The lower the p-value, the stronger the evidence that your variant truly outperforms the control.

95% vs 99% confidence — when should I use which?

Use 95% confidence for routine UI and copy tests where the cost of being wrong is low. Use 99% for changes that affect pricing, checkout flow, or anything you can't reverse easily — the higher bar reduces the chance of shipping a false positive on a high-stakes change.

Why does my test "go non-significant" the longer I run it?

Reverting confidence usually means the early lift was noise, not signal. Small samples produce volatile conversion rates. Trust the longer-running result — and avoid stopping a test the moment it crosses significance, which inflates false-positive rates substantially.

What sample size do I need?

Sample size depends on baseline conversion rate and minimum detectable effect (MDE). Tools like Optimizely's calculator give you a pre-test number. This calculator shows post-hoc significance — what your current data tells you after the test has run.

How does Pace use A/B test data?

Pace doesn't run on-site A/B tests — but it tracks ad-creative performance and surfaces fatigued creatives (Meta: frequency above 3 with CTR below 1%) so you know exactly when to rotate. The same statistical thinking that powers this calculator informs how Pace's Creative Lens flags ads.

Want creative rotation tied to performance — not your gut feel?

Pace tracks creative-level CTR, CPA, and ROAS across Meta, TikTok, and Google — and surfaces creatives showing fatigue or runaway lift before the data goes stale.

This calculatorWith Pace
FrequencyOne-shot, when you rememberDaily significance checks
ChannelsOne test at a timeMeta, TikTok, and Google creatives
ActionTells you the p-valueSurfaces winners and laggards
HistoryResets on refreshRolling creative-fatigue trail

Ready to stay on pace?

14-day free trial on the Enterprise plan.