🙌 Our latest webinar is live!
Sep 21
Incrementality

Incrementality Testing: How to Run Holdout and Geo Tests

Andre Sottil
Founder & CEO, Admira

Incrementality testing measures how many conversions your marketing actually caused, not just touched. You compare a group exposed to your ads against a statistically similar group that was held out, and the difference in outcomes is your true lift. The two most practical methods are audience holdout tests, where a random slice of users sees no ads, and geo tests, where you change spend in matched regions. Both work without cookies or user-level tracking, which makes them the most privacy-resilient measurement tools available today.

Why attribution alone cannot answer this question

Attribution models, whether last click or data driven, distribute credit among the touchpoints that preceded a conversion. They cannot tell you whether the conversion would have happened anyway. Retargeting is the classic example: it reaches people who are already close to buying, so it earns strong attribution numbers while often adding little incremental revenue.

Incrementality testing fixes this by building a counterfactual. The holdout group shows you what happens without the ad, so the comparison isolates cause from correlation. It is the logic of a clinical trial applied to media spend, and it is the only way to separate demand your ads created from demand they merely intercepted.

How to run an audience holdout test

An audience holdout test randomly splits eligible users into a test group that sees your ads and a control group that does not. Because the split is random, any difference in conversion rate between the groups is attributable to the ads themselves.

  1. Pick one channel and one conversion event. Testing everything at once produces an unreadable result.
  2. Split the audience randomly, typically 80 to 90 percent test and 10 to 20 percent control.
  3. Suppress ads for the control group using the platform’s lift study tooling, since you usually cannot enforce this yourself inside walled gardens.
  4. Run until each group has at least a few hundred conversions. Thin data produces confidence intervals too wide to act on.
  5. Compare conversion rates and calculate lift: test rate minus control rate, divided by control rate.

The main limitations are that you depend on the platform’s own tooling and honesty, and that cross-device behavior can contaminate the control group when a held-out user simply sees the ad on another device.

How to run a geo holdout test

A geo test uses regions instead of users, which means no tracking of individuals at all. You change spend in some markets and hold it steady in comparable ones, then read the difference in aggregate sales.

  1. Select markets that behave similarly on historical sales, seasonality, and population, so the comparison starts from a fair baseline.
  2. Assign treatment and control markets, ideally with a synthetic control method that builds a weighted comparison from several regions.
  3. Make a meaningful spend change. Going dark in treatment markets, or raising spend 50 to 100 percent, produces a readable signal; a 10 percent tweak will drown in noise.
  4. Run four to eight weeks, then compare actual sales in treatment markets against the modeled counterfactual.

Watch for contamination: national promotions, press coverage, or delivery zones that cross market borders can blur the comparison and hand you a lift number you cannot trust.

Holdout vs geo: which to choose

DimensionAudience holdoutGeo test
Best forSingle-platform questions (Meta, Google)Channel-level or cross-channel questions
Tracking requiredPlatform user-level dataNone, only regional sales
Typical duration2 to 4 weeks4 to 8 weeks
Privacy resilienceModerate, depends on platformHigh, fully aggregate
Main riskContaminated control groupPoorly matched markets

Use audience holdouts when you want a fast read on one platform’s performance. Use geo tests when you want a platform-independent answer, when tracking is limited, or when the channel has no reliable click signal, like TV, audio, or influencer campaigns.

Mistakes that invalidate tests

The most common failure is running a test that is too small or too short, then acting on noise as if it were signal. The second is testing during promotional peaks like Black Friday, when demand spikes swamp the media effect you were trying to isolate. Others include stopping early because the result looks good, changing creative or budgets mid-test, and treating a single result as permanent truth when incrementality drifts as audiences saturate over time.

Reading the result without fooling yourself

A lift number arrives with a confidence interval, and ignoring that range is how good tests produce bad decisions. If your test says a channel drove 12 percent incremental lift but the interval runs from minus 3 to plus 27 percent, you have not proven the channel works; you have proven your test was underpowered. Treat a wide interval as a signal to run longer or concentrate spend, not as license to round up to the headline figure. It is also worth translating lift into incremental ROAS and comparing it against your margin, because a channel can post a statistically real lift that is still unprofitable once you price in the media it took to produce. The point of testing is a decision you can defend, and a defensible decision respects the uncertainty the math actually carries.

Where testing fits your everyday measurement

Incrementality tests are periodic and precise, but you cannot run them on every channel every week. The practical model is to use lift tests to calibrate the attribution and marketing mix modeling you rely on daily, so your everyday numbers inherit the truth the tests reveal. This is the gap Admira is built to close: it unifies multi-touch attribution, marketing mix modeling, and incrementality testing on a single cookieless-first foundation, so a holdout or geo result does not end its life in a slide deck but flows straight back into the dashboards your team reads every morning. If you are tired of running lift tests that never quite reconcile with your day-to-day reporting, book a demo and we will map your channels to the tests worth running first.

FAQ

How long should an incrementality test run?

Audience holdouts usually need two to four weeks to accumulate enough conversions for a readable result. Geo tests typically need four to eight weeks because regional sales data is noisier and you need both a clean pre-period and post-period to compare against. Running shorter than that is how teams end up acting on noise.

Can smaller brands afford incrementality testing?

Yes, with adjustments. Platform lift studies are often free above a modest spend threshold, and a simple two-market geo test costs only the media you were already buying. The real constraint is conversion volume, not budget, so smaller brands should test their single largest channel first rather than measuring everything at once.

What is a good incremental ROAS?

There is no universal benchmark. Incremental ROAS is usually lower than platform-reported ROAS, sometimes dramatically so for retargeting, because platforms take credit for conversions that would have happened anyway. The useful comparison is against your margin threshold: if a channel is incrementally profitable, scale it; if not, reallocate.

How often should I repeat incrementality tests?

Re-test your biggest channels once or twice a year, and re-test after major changes like a new creative strategy, an audience expansion, or a significant budget shift. Incrementality drifts as audiences saturate, so a result is a snapshot, not a permanent truth. Use each test to recalibrate your everyday attribution.