Why the number on the dashboard is already too high
Every A/B test is a comparison of noisy measurements. Even a variant with zero true effect will occasionally show a positive lift purely from the randomness of which visitors happened to land in which bucket that week. Run the test long enough, or run enough variants, and some of what gets measured as a 'winner' is really just noise that ran in the winning direction.
The problem is structural, not a mistake anyone made: conditional on a variant having been selected as the winner, it is statistically more likely to have benefited from favorable noise than a variant picked at random would be. This is the same phenomenon economists call the winner's curse in auctions — the winning bidder is, on average, the one who most overestimated the item's value. In testing, the 'winning' variant is, on average, the one whose test-window noise most overestimated its true effect.
Peeking at results daily and stopping the moment a variant crosses significance makes this worse, not better. That stopping rule maximizes the chance of catching a variant exactly when its cumulative noise happens to be at a favorable peak — which is precisely the moment a test dashboard looks most convincing and is least representative of the long-run truth.
A test reporting +9% at the exact day it crosses significance is disproportionately likely to be reporting a peak, not a plateau.
What the gap actually looks like in practice
The pattern is recognizable to anyone who has run CRO programs for more than a year: a checkout or pricing-page test reports a clean, significant lift, the team ships it immediately, and the metric in production settles measurably below what the test promised within a month or two. Nobody can point to a specific error. The change did work — it's just that part of the reported number was never real to begin with.
Published analysis and practitioner write-ups on this exact pattern converge on a similar magnitude: reported lifts typically overstate realized, in-production lifts by somewhere in the range of 20% to 50%. A test claiming +9% settling at roughly +4-7% once measured cleanly post-launch is well within that range, not an anomaly.
This matters most for agencies and in-house teams reporting results upward. A quarter of 'wins' each carrying an unacknowledged 20-50% inflation compounds into a growth narrative that drifts further from actual revenue every quarter it goes unchecked — and the gap eventually surfaces as an executive asking why the roadmap of wins hasn't produced the revenue those wins implied.
The fix isn't a stricter p-value
The instinctive fix — require a lower p-value or a longer minimum runtime — helps only marginally, because the underlying mechanism (selecting a winner from noisy data) doesn't go away just because the bar is higher. A stricter threshold reduces how often you declare a false winner; it does very little to correct the size of the lift you report once you have declared one.
The actual fix has two parts, and both are cheap relative to what a single inflated result costs when it's used to justify a roadmap.
Set the sample size before the test starts, and don't stop early
Calculate the sample size a real effect of your minimum meaningful size would need, commit to it before launch, and let the test run to that number regardless of what the interim dashboard shows. This alone removes the worst of the early-stopping inflation.
Re-validate with a post-launch holdout
Keep a small slice of traffic on the old experience for four to eight weeks after shipping the 'winner,' and compare it against the new experience over that window. This measures the actual, settled effect of what shipped — not the effect measured during the test's own decision window, which is the number most prone to the winner's curse.
Report the holdout-corrected number, not the test-tool number
The figure your testing platform shows at the moment of decision is an estimate, not a guarantee. When reporting results to a client or an executive team, either wait for the holdout re-validation or explicitly flag the initial number as provisional and likely to shrink.
What to do with a year of past 'wins'
You don't need to re-run every historical test to get value from this. Pull the handful of tests from the last year with the largest reported lifts — the winner's curse effect scales with how extreme a result is, so the biggest reported wins are the ones most worth checking against actual downstream metrics.
Compare the metric's trajectory in the months after each of those launches against the lift the test originally reported. Tests where the metric has genuinely held near the reported number are the real wins. Tests where it settled well below are candidates for the holdout-validation discipline going forward, and worth a quiet correction in whatever roadmap or client report originally cited the bigger number.
The goal isn't to distrust every past result. It's to build the habit of treating a test's reported lift as a starting estimate that gets confirmed or revised by what actually happens after launch — the same discipline that makes any of this site's other measurement work trustworthy in the first place.
Terms used in this piece
We do this for a living
If you'd rather not build this yourself, these are the services where it lives.