Run enough variants, or run one variant long enough through naturally noisy traffic, and some of what gets measured as "lift" is just the test period's luck landing in the winning variant's favor. The stopping rule most teams use — declare a winner the moment it crosses significance — is exactly the rule most likely to stop at a lucky peak rather than the true effect.
This isn't a data-quality problem or a tracking bug. It's a structural property of selecting a winner from noisy data: whichever variant you pick, conditional on it having won, is more likely to have benefited from favorable noise than any variant chosen at random would be. The bigger the pool of variants or metrics tested, the larger this inflation tends to run.
The practical result shows up as a familiar pattern: a test reports a confident, statistically significant lift, the team ships it, and the metric in production settles noticeably lower than what the test dashboard promised — not because the change did nothing, but because part of the reported number was never real to begin with.
Why Winner's Curse matters
Every roadmap and every client report that cites a test's raw reported lift as the ongoing, permanent effect is very likely reporting an inflated number — and compounding that error across a year of "wins" produces a growth story with a widening gap to actual revenue.
A checkout test with a real, smaller effect
A checkout variant tests at +9% conversion with a p-value that clears significance the day the test hits its planned sample size, and the team ships it immediately. A holdout re-validation over the following month — a slice of traffic still on the old checkout, watched for eight weeks after launch — shows the true, stable lift is closer to +4%. The other five points were the winner's curse: real variance in that test window that happened to favor the new checkout, reported as if it were permanent.
Benchmarks
- Typical gap between reported and realized lift
- 20–50% overstatement
- Recommended post-launch holdout window
- 4–8 weeks
Ranges drawn from Digital Squad client accounts and published industry data. Treat them as orientation, not targets — your category may differ substantially.
Common mistakes
Stopping the moment significance is crossed
Peeking at results daily and stopping the instant a variant clears the threshold inflates the winner's curse further, because it maximizes the chance of catching a variant at a lucky peak rather than at its true, settled effect.
Reporting the test's number as the shipped number
The figure in the testing tool is the estimate at the moment of the decision, not a guarantee of what production will show. Reporting it to a client or exec team without a discount or a re-validation plan sets an expectation the metric usually won't meet.
Skipping post-launch holdout validation
Keeping a small slice of traffic on the old experience for a few weeks after shipping the "winner" is the cheapest way to learn the true effect size — and most teams skip it because the test already declared a winner, which feels like the work is done.
Where we work on this