Holdout Testing vs Incrementality Testing: Standing Check or One-Off Study
In short: Both use a randomized control to isolate the ad effect from what would have happened anyway. Holdout testing is a specific technique: you carve off a slice of one audience, usually CRM or retargeting, suppress ads to it indefinitely, and watch the gap over time. Incrementality testing is the broader practice of running a designed, time-boxed experiment - which can use an audience holdout, a geo split, or a platform's own lift tool - to answer one specific budget question. If you want a standing sentinel on your riskiest audience, run a holdout; if you need a decision-grade answer for one specific question, design a proper incrementality test.
By the AdFlint research team · Fact-checked against current Google and Meta platform behavior · Last reviewed July 2026
Holdout Testing
Permanently excludes a randomly selected share of your audience from seeing ads, so their behavior serves as an ongoing baseline.
A fixed percentage of users or accounts is suppressed from targeting for the duration of the test, and the gap between held-out and exposed groups gives you the incremental effect. Advertisers keep a standing holdout on retargeting and CRM audiences, where reported returns are most likely to be double-counted demand. Watch for leakage: if the held-out group can still be reached by other campaigns, the baseline is contaminated.
Full definitionIncrementality Testing
Measures how many conversions the advertising actually caused by comparing an exposed group against a randomized control that saw nothing.
You withhold ads from a randomly chosen slice of the eligible audience or market, then read the difference in outcomes between exposed and held-out groups. It answers the question attribution cannot: how much of this would have happened anyway. Advertisers run it on brand search, retargeting, and any channel with suspiciously good reported returns. The mistake is running it too small or too briefly to detect a realistic lift.
Full definitionSide by side.
The differences that actually change what happens in your account.
| Holdout Testing | Incrementality Testing | |
|---|---|---|
| Duration | Indefinite - the suppressed slice stays out of targeting for as long as the campaign runs. | Time-boxed - a defined start and end date tied to one decision. |
| What gets randomized | A slice of one audience list, commonly CRM or retargeting. | Whatever unit fits the question - an audience split, a geo split, or a platform lift test. |
| Why you run it | An ongoing sanity check that a specific audience is not just capturing demand that would have converted anyway. | A specific budget or strategy decision that needs a causal answer before money moves. |
| Setup effort | Low - build an exclusion list or audience split and leave it running. | Higher - requires a pre-registered hypothesis, a sample size estimate, and a defined read window. |
| Where the result lives | A running comparison you re-check periodically, not a single report. | A single readout delivered at the end of the test window. |
| Failure mode | Leakage - other campaigns, organic, or email still reach the held-out group and erase the gap. | An underpowered test - too short or too small a sample to detect a realistic lift, so you conclude 'no effect' when you just lacked the sample. |
| Typical target | Retargeting and CRM audiences, where reported returns are most likely to be inflated. | Any single channel or the whole marketing budget, depending on what the test is scoped to answer. |
What actually separates them.
Holdout testing keeps running until you turn it off; incrementality testing has a planned end date because it is built to answer one question and close.
Holdout testing only ever splits an audience list; incrementality testing can split by audience, by geography, or run through a platform's built-in lift tool depending on what the channel supports.
A holdout requires no statistical planning beyond picking a split size; a proper incrementality test needs a minimum detectable effect and sample size worked out before it starts, or the null result is meaningless.
Holdout testing is almost always self-managed inside your own audience tooling; incrementality testing frequently borrows platform infrastructure or a third-party geo-testing tool.
Because a holdout never ends, you are trading a small, permanent tax on delivery for a rolling read; a discrete incrementality test trades a temporary, sharper disruption for a one-time definitive answer.
Which one should you use?
Use Holdout Testing when
- Your retargeting or CRM campaigns report suspiciously strong returns and you want an always-on check rather than a one-time verdict.
- You want a low-effort, low-maintenance signal that does not require designing a formal experiment.
- You are comfortable permanently sacrificing a small slice of reach in exchange for a continuous baseline.
- You manage the audience lists directly and can enforce the suppression without relying on a platform's test product.
Use Incrementality Testing when
- You have a specific decision on the table, like whether to cut or grow a channel's budget, and need a defensible answer.
- The question spans more than one audience or channel, so a simple list-based holdout will not cover it.
- You can tolerate a temporary, larger disruption to delivery in exchange for a sharper, time-boxed result.
- You need to size the test properly first, because a vague or underpowered read will not hold up to scrutiny.
Common questions.
Is holdout testing a type of incrementality testing?
Yes. Holdout testing is one way to run an incrementality test - it is the standing, audience-based version of it. Geo testing and platform conversion lift tests are other ways to do the same underlying thing: compare an exposed group against a randomized group that saw nothing.
How big should my holdout group be?
There is no universal number - it depends on how much reach you can afford to sacrifice and how big a gap you need to detect above normal week-to-week noise. Start with a modest slice, watch whether the gap holds steady across a few cycles, and adjust the size based on how clean or noisy the read looks, rather than chasing a fixed figure.
Can I run a holdout and a separate incrementality test at the same time?
Yes, but check they are not drawing from the same pool, or you will contaminate both reads. A standing holdout on your CRM list and a discrete geo test on a different channel can run in parallel without interfering with each other.
Why did my holdout test show no difference?
Either the audience genuinely was not incremental, or the group leaked - other campaigns, organic search, or email kept touching the supposedly excluded group and closed the gap. Check for leakage across every channel that can reach that audience before concluding the spend added nothing.
Or stop choosing between them.
AdFlint picks the setting, writes the ads, and keeps optimizing inside the Google and Meta accounts you already own.
Related comparisons
- A/B Testing vs Incrementality Testing
- A/B Testing vs Conversion Lift Tests
- A/B Testing vs Brand Lift Tests
- A/B Testing vs Geo Testing
- A/B Testing vs Holdout Testing
- A/B Testing vs Marketing Mix Modeling
- A/B Testing vs Multi-Touch Attribution
- Conversion Lift Tests vs Incrementality Testing