Most holdout tests fail before they start
Underpowered or too short, they report no effect and everyone treats that as an answer. Two numbers decide it, and both are knowable in advance.
Your numbers
Metros or regions you can target separately.
Where the channel is switched off.
Across all markets, not just the holdout.
Coefficient of variation in weekly revenue. 15–30% is typical.
First touch to purchase. The test must outlast it.
How much you believe this channel contributes.
Smallest effect this design can detect
7.1%
10 holdout vs 30 control markets over 10 weeks, 95% confidence.
Revenue withheld
$120,000
If the channel really contributes 12%. If it contributes nothing, this costs nothing.
Minimum duration
9 wk
1.5× your purchase cycle, so the effect has time to appear.
Weeks needed to detect your expected effect
9 wk
Your planned duration is sufficient.
Sound design: 7.1% detectable against an expected 12%, over a period that outlasts the purchase cycle. Withheld revenue is around $120,000 if the channel performs as you believe. Write down what the result will change before you start — a holdout that produces an uncomfortable answer will otherwise get reinterpreted.
Calculated in your browser — nothing is sent anywhere, and nothing is stored.
The maths, so it isn't a black box.
It computes the smallest effect you could detect
From the number of holdout and control markets, the run length, and how much your weekly revenue naturally varies. If that number is larger than the effect you expect to find, the test cannot answer your question no matter how carefully you run it.
It enforces a duration against your purchase cycle
A test shorter than the buying cycle measures the delay rather than the effect — people exposed before the switch-off are still converting during it. We require one and a half times the cycle as a floor.
It prices the revenue you withhold
Holdout markets' share of revenue, across the run, multiplied by the effect you believe the channel has. The cost is proportional to how well the channel actually works, which means a test on a channel doing nothing costs nothing.
And it solves for the duration you actually need
If the design is underpowered, the useful output is not a warning but a number: how many weeks would make it answerable at this split.
Questions about this calculation
A two-sample difference in means across market-weeks, at 95% confidence and 80% power. Each market-week is treated as an observation, so the detectable effect falls with more markets, more weeks and lower week-to-week variance. It is deliberately a simplification — it ignores the serial correlation real geographies carry, which means a proper difference-in-differences analysis will do slightly better than this predicts. Erring conservative is the right direction for planning, because an underpowered test that reports 'no effect' is the expensive failure.
Terms used here
Other calculators
Where we do this work
We'll build this model on your real data.
A free 30-minute teardown where we build your contribution margin model from actual figures and show you where the funnel leaks. Yours to keep either way.