Answer first
What this calculator tells you
Test whether the gap between two conversion rates is likely to be real, with the z-score and p-value. Decide whether a test result is worth acting on or could easily be chance. Formula: z = (p_B − p_A) ÷ √(p̂(1 − p̂)(1/n_A + 1/n_B)), with p̂ the pooled rate; two-tailed p = 2 × (1 − Φ(|z|)). At the worked-example inputs, the relative uplift of b over a is 20.0%. Holding every other input steady, moving visitors, version a from 4,000 to 6,000 moves the result from -4.0% to 44.0%.
Transparent method
The formula
Decide whether a test result is worth acting on or could easily be chance.
Worked example
Example inputs
How to interpret the result
An A/B test compares two conversion rates, and the question is whether the gap is bigger than chance would produce. With 250 conversions from 5,000 visitors against 300 from 5,000, version B is up 20 percent, the z-score is about 2.19, and the p-value is about 0.028. That is below 0.05, the usual bar, so the gap is unlikely to be luck alone, though it is not proof.
At the worked-example inputs the relative uplift of b over a is 20.0%. It rises with conversions, version b and visitors, version a and falls as visitors, version b and conversions, version a increase.
These are planning metrics, not audited accounting or a valuation opinion.
Before you rely on it
What to check
Decide the sample size and the stopping rule before you start. Stopping the moment a result looks good raises the false positive rate well above the stated p-value.
The common error
Where people go wrong with a/b test significance calculator
Reading a p-value of 0.028 as a 97 percent chance that B is better. It is the chance of a gap this large if there were no real difference, which is a different statement.
Sensitivity evidence
How visitors, version a changes the relative uplift of b over a
Holding every other input at the worked-example value, moving visitors, version a from 4,000 to 6,000 moves the relative uplift of b over a from -4.0% to 44.0%: a spread of 48.0%, or 240% of the worked-example result.
| Visitors, version A | Relative uplift of B over A | Z-score | Two-tailed p-value |
|---|---|---|---|
| 4,000 | -4.0% | -0.492 | 0.623 |
| 4,500 | 8.0% | 0.926 | 0.354 |
| 5,000worked example | 20.0% | 2.2 | 0.028 |
| 5,500 | 32.0% | 3.3 | 0.00083 |
| 6,000 | 44.0% | 4.4 | 0.00001 |
Every input, tested
Which input moves the relative uplift of b over a most
Of the 4 inputs, conversions, version b moves the relative uplift of b over a most (24.0% across the range tested) and conversions, version a moves it least (24.2%).
| Input | Tested from | To | Relative uplift of B over A at each end | Swing |
|---|---|---|---|---|
| Conversions, version B | 270 | 330 | 8.0% to 32.0% | 24.0% (120%) |
| Visitors, version A | 4,500 | 5,500 | 8.0% to 32.0% | 24.0% (120%) |
| Visitors, version B | 4,500 | 5,500 | 33.3% to 9.1% | 24.2% (121%) |
| Conversions, version A | 225 | 275 | 33.3% to 9.1% | 24.2% (121%) |
Two variables at once
Relative uplift of B over A by visitors, version a and conversions, version a
Across the grid the relative uplift of b over a runs from -20.0% to 80.0%. Moving visitors, version a from 4,000 to 6,000 shifts it by 48.0% at the middle column, and moving conversions, version a from 200 to 300 shifts it by 50.0% at the middle row, so conversions, version a is the bigger lever here.
| Visitors, version A \ Conversions, version A | 200 | 250 | 300 |
|---|---|---|---|
| 4,000 | 20.0% | -4.0% | -20.0% |
| 4,500 | 35.0% | 8.0% | -10.0% |
| 5,000 | 50.0% | 20.0% | 0.000% |
| 5,500 | 65.0% | 32.0% | 10.0% |
| 6,000 | 80.0% | 44.0% | 20.0% |
The highlighted cell is the worked example: 20.0%.
Step by step
The worked example, input by input
| Input | Value used | What it means |
|---|---|---|
| Visitors, version A | 5,000 | Enter the visitors, version a used in this calculation. |
| Conversions, version A | 250 | Enter the conversions, version a used in this calculation. |
| Visitors, version B | 5,000 | Enter the visitors, version b used in this calculation. |
| Conversions, version B | 300 | Enter the conversions, version b used in this calculation. |
| Relative uplift of B over A | 20.0% | |
| Z-score | 2.2 | |
| Two-tailed p-value | 0.028 | |
Inputs, definitions and assumptions
Visitors, version A
Enter the visitors, version a used in this calculation. The prefilled worked-example value is 5,000.
Conversions, version A
Enter the conversions, version a used in this calculation. The prefilled worked-example value is 250.
Visitors, version B
Enter the visitors, version b used in this calculation. The prefilled worked-example value is 5,000.
Conversions, version B
Enter the conversions, version b used in this calculation. The prefilled worked-example value is 300.
How to use this calculator
- 1Verify the inputs. Gather visitors, version a, conversions, version a, visitors, version b and conversions, version b from your own documents; the prefilled values are examples.
- 2Save a baseline. The worked example puts the relative uplift of b over a at 20.0%. Store your own version of it as Scenario A.
- 3Test one change. Start with conversions, version b, the input with the biggest effect here: moving conversions, version b from 270 to 330 takes the relative uplift of b over a from 8.0% to 32.0%, a swing of 120% of the worked-example figure.
- 4Check the boundary. Read the interpretation boundary above before acting on the result.
People also ask
Frequently asked questions
How do you calculate a/b test significance?
z = (p_B − p_A) ÷ √(p̂(1 − p̂)(1/n_A + 1/n_B)), with p̂ the pooled rate; two-tailed p = 2 × (1 − Φ(|z|)). At the worked-example inputs the relative uplift of b over a is 20.0%.
What does the a/b test significance result mean?
Decide whether a test result is worth acting on or could easily be chance. At the worked-example inputs the relative uplift of b over a is 20.0%. It rises with conversions, version b and visitors, version a and falls as visitors, version b and conversions, version a increase.
How much does visitors, version a change the relative uplift of b over a?
Holding every other input at the worked-example value, moving visitors, version a from 4,000 to 6,000 moves the relative uplift of b over a from -4.0% to 44.0%, a spread of 48.0%.
What are the limits of this a/b test significance calculator?
These are planning metrics, not audited accounting or a valuation opinion. The tables on this page test visitors, version a only from 4,000 to 6,000; a value outside that range is not tabulated here.
Which input moves the relative uplift of b over a most in the a/b test significance calculator?
Ranked by how far each moves the relative uplift of b over a across the range tested: conversions, version b (24.0%, 120%), visitors, version a (24.0%, 120%), visitors, version b (24.2%, 121%) and conversions, version a (24.2%, 121%).
How much does conversions, version a matter in the a/b test significance calculator?
The worked example uses 250. Holding every other input at its worked-example value, moving conversions, version a from 225 to 275 takes the relative uplift of b over a from 33.3% to 9.1%, a swing of 121% of the worked-example figure.
How much does visitors, version b matter in the a/b test significance calculator?
The worked example uses 5,000. Holding every other input at its worked-example value, moving visitors, version b from 4,500 to 5,500 takes the relative uplift of b over a from 33.3% to 9.1%, a swing of 121% of the worked-example figure.
How much does conversions, version b matter in the a/b test significance calculator?
The worked example uses 300. With the other inputs left at the worked example, moving conversions, version b from 270 to 330 takes the relative uplift of b over a from 8.0% to 32.0%, a swing of 120% of the worked-example figure.
Which inputs change the z-score in the a/b test significance calculator?
At the worked-example inputs it is 2.2. Visitors, version a takes it from 0.926 to 3.3, conversions, version a takes it from 3.4 to 1.1, visitors, version b takes it from 3.5 to 1 and conversions, version b takes it from 0.901 to 3.4.
Which inputs change the two-tailed p-value in the a/b test significance calculator?
At the worked-example inputs it is 0.028. Visitors, version a takes it from 0.354 to 0.00083, conversions, version a takes it from 0.00077 to 0.283, visitors, version b takes it from 0.00051 to 0.296 and conversions, version b takes it from 0.368 to 0.00062.
How do I price a new product or service?
Start from the margin you need instead of from a markup on cost, because the two are computed on different bases and confusing them consistently underprices. Then check the price against what the market will bear and against your break-even volume. A price that is theoretically correct and commercially unsellable is not a price.
Sources and evidence
Free Calculators Online is independent and is not affiliated with or endorsed by the source organizations. Educational estimates only.