Why the two measures disagree
Attribution is observational. It records which touchpoints appeared in a converting journey and applies a rule to divide credit. It cannot distinguish between a touchpoint that caused a conversion and one that merely happened to be present.
Brand search is the clearest illustration. Someone types your company name, sees your ad, clicks, and buys. Attribution credits the ad with the full conversion. But that person was looking for you specifically — in most cases they would have clicked the organic result immediately below at no cost.
This isn't a flaw in the platforms; they are reporting what they observe. It is a limitation of observation. Establishing causation requires withholding the treatment from someone, which is exactly what an incrementality test does.
The geo holdout, step by step
A geo holdout splits your market into regions that receive advertising and regions that don't, then compares total outcomes. It is the most practical incrementality method for most businesses.
Pick comparable markets
Test and control regions need similar historical revenue trends, seasonality, and customer mix. Use 12 months of history to verify they track each other before the test — mismatched markets produce confident wrong answers.
Withhold entirely in control
Reducing spend in control regions rather than stopping muddies the result. Partial treatment produces a partial signal that is hard to interpret.
Run 4–6 weeks minimum
Long enough to cover multiple purchase cycles and let delayed conversions land. Two-week tests measure mostly noise in almost every category.
Measure total revenue, not attributed
The comparison is total revenue in test markets versus total in control, from your own systems. Using platform-attributed data reintroduces exactly the bias you are testing for.
Account for the post-test period
Watch control markets for a week or two after resuming. A rebound suggests demand was delayed rather than lost, which changes the interpretation.
Reading the result
Incremental lift is the difference in total revenue between test and control markets, normalized for their historical relationship. Dividing incremental revenue by spend gives you incremental ROAS.
The number will usually be lower than your attributed ROAS, often substantially. A channel showing 4.2x attributed and 1.6x incremental is telling you that most of what it claims credit for was going to happen regardless.
That is not automatically a reason to cut it. A channel with modest incrementality but a genuinely low absolute cost may still be worth running. What the test gives you is the ability to make that decision on real numbers rather than a dashboard that structurally overstates.
Test brand search first. It is the easiest to run, the most commonly over-credited, and the result most likely to change how you allocate budget.
When you can't run a geo test
Geo holdouts need enough geographic spread and volume. Smaller or single-market businesses have other options, each with tradeoffs.
Platform lift studies
Meta and Google both offer built-in conversion lift tests using a randomized holdout. Convenient, but you are asking the platform to grade its own work — treat directionally.
Time-based on and off tests
Pausing a channel for a defined period and comparing. Weaker, because seasonality and external events confound it. Alternate on and off periods several times to reduce that.
Post-purchase attribution surveys
One question at checkout. Self-reported and imperfect, but independent of pixel tracking — so when it disagrees sharply with your platforms, the direction of disagreement is informative.
Marketing mix modelling
Statistical modelling across channels and time. Powerful at scale, but it needs two to three years of history and meaningful spend variation to be reliable.
Make it routine, not a one-off
A single incrementality test tells you about one channel at one moment. Markets shift, competitors enter, creative fatigues, and a channel that was incremental last year may not be now.
Build a testing calendar: one channel per quarter, starting with whatever consumes the most budget. Over two years you accumulate a genuine map of what is actually driving growth.
The organizational benefit matters as much as the measurement. Once a team has seen attributed and incremental numbers diverge on their own account, the conversation about budget allocation changes permanently.
Weighing the decision?
Terms used in this piece
We do this for a living
If you'd rather not build this yourself, these are the services where it lives.