Statistical significance
Statistical significance is the verdict that a difference seen in a test is unlikely to be chance alone at a threshold set before the test began.
How it works
Even two equally good versions of a page usually show different conversion rates: orders arrive by chance. The p-value is how often equal versions would show a gap at least as large as yours. Below a level fixed in advance, commonly 5% (a 95% confidence level), the difference counts as significant.
Significance says the gap is probably not chance, not how much it matters. Stopping an A/B test at the first significant reading produces false winners far more often than the stated level, so fix the sample size in advance. The A/B testing guide for small stores covers what to do when traffic is short.
Formula
z = (CR of B − CR of A) ÷ √(p × (1 − p) × (1/n of A + 1/n of B)); CR is conversion rate, n clicks, p the pooled rate of both versions; at 5%, z beyond ±1.96 is significant
Example
Example store, not client data.
The tableware shop splits a month’s 8,000 clicks between two layouts. A: 150 orders from 4,000 clicks, 3.75%. B: 170 from 4,000, 4.25%, a 13% lift. Pooled: 4%. z = 0.005 ÷ √(0.04 × 0.96 × 0.0005) = 1.14, p-value ≈ 0.25: not significant. Detecting this lift at 5% with 80% power takes about 23,100 clicks per version: almost six months’ traffic.
How GetProfit reads it
Instead of a significance test, the portal’s change feed judges a change by fixed bands over 14 days: worked if ROAS rises above 1.2× with conversions at 0.8× or more; harmed if ROAS falls below 0.8× or conversions below half; no verdict under 7 days or with no conversions. For scale, in GetProfit data a store’s monthly ROAS typically deviates 16% from its own median (median of 110 stores with at least 8 months of data, June 2025 – June 2026).
Right and wrong readings
- Wrong: “Significant at 95%, so B is better with 95% probability.” Right: 5% is how often equal versions would still show a significant gap, not the chance that B wins.
- Wrong: “Not significant, so the layout changes nothing.” Right: the example lacked traffic; the 13% lift may be real.
Sources
- 7.1.3.1. Critical values and p values (NIST/SEMATECH e-Handbook of Statistical Methods) — significance level, set in advance. Checked 2 October 2026.
- How Not To Run an A/B Test (Evan Miller) — peeking at results inflates false positives. Checked 2 October 2026.
- Sample Size Calculator (Evan Miller) — sample size per version (baseline 3.75%, effect 0.5 points absolute, power 80%, α 5%). Checked 2 October 2026.
- GetProfit data: 110 online stores, June 2025 – June 2026 — monthly ROAS deviation from each store’s median.