Skip to content
Free tool · No signup

Most holdout tests fail before they start

Underpowered or too short, they report no effect and everyone treats that as an answer. Two numbers decide it, and both are knowable in advance.

Your numbers

Metros or regions you can target separately.

Where the channel is switched off.

$

Across all markets, not just the holdout.

%

Coefficient of variation in weekly revenue. 15–30% is typical.

wk
wk

First touch to purchase. The test must outlast it.

%

How much you believe this channel contributes.

Smallest effect this design can detect

7.1%

10 holdout vs 30 control markets over 10 weeks, 95% confidence.

Revenue withheld

$120,000

If the channel really contributes 12%. If it contributes nothing, this costs nothing.

Minimum duration

9 wk

1.5× your purchase cycle, so the effect has time to appear.

Weeks needed to detect your expected effect

9 wk

Your planned duration is sufficient.

Sound design: 7.1% detectable against an expected 12%, over a period that outlasts the purchase cycle. Withheld revenue is around $120,000 if the channel performs as you believe. Write down what the result will change before you start — a holdout that produces an uncomfortable answer will otherwise get reinterpreted.

Calculated in your browser — nothing is sent anywhere, and nothing is stored.

How it works

The maths, so it isn't a black box.

01

It computes the smallest effect you could detect

From the number of holdout and control markets, the run length, and how much your weekly revenue naturally varies. If that number is larger than the effect you expect to find, the test cannot answer your question no matter how carefully you run it.

02

It enforces a duration against your purchase cycle

A test shorter than the buying cycle measures the delay rather than the effect — people exposed before the switch-off are still converting during it. We require one and a half times the cycle as a floor.

03

It prices the revenue you withhold

Holdout markets' share of revenue, across the run, multiplied by the effect you believe the channel has. The cost is proportional to how well the channel actually works, which means a test on a channel doing nothing costs nothing.

04

And it solves for the duration you actually need

If the design is underpowered, the useful output is not a warning but a number: how many weeks would make it answerable at this split.

Questions about this calculation

A two-sample difference in means across market-weeks, at 95% confidence and 80% power. Each market-week is treated as an observation, so the detectable effect falls with more markets, more weeks and lower week-to-week variance. It is deliberately a simplification — it ignores the serial correlation real geographies carry, which means a proper difference-in-differences analysis will do slightly better than this predicts. Erring conservative is the right direction for planning, because an underpowered test that reports 'no effect' is the expensive failure.

Want us to run it properly

We'll build this model on your real data.

A free 30-minute teardown where we build your contribution margin model from actual figures and show you where the funnel leaks. Yours to keep either way.