Holdout Testing vs Multi-Touch Attribution: Suppress or Stitch
In short: Both work at the individual level, but one manipulates exposure and the other only observes it. Holdout testing randomly blocks a slice of a named audience from seeing ads and reads the gap against everyone else as the true incremental effect. Multi-touch attribution stitches together the touchpoints on paths that did convert and splits credit among them, without ever knowing what would have happened if a touch were removed. Holdout testing answers whether an audience is incremental at all; MTA answers which touchpoints showed up on the paths that converted. If you need a trustworthy yes-or-no on one audience, run a holdout; if you need fast, granular, directional signal across campaigns, use MTA.
By the AdFlint research team · Fact-checked against current Google and Meta platform behavior · Last reviewed July 2026
Holdout Testing
Permanently excludes a randomly selected share of your audience from seeing ads, so their behavior serves as an ongoing baseline.
A fixed percentage of users or accounts is suppressed from targeting for the duration of the test, and the gap between held-out and exposed groups gives you the incremental effect. Advertisers keep a standing holdout on retargeting and CRM audiences, where reported returns are most likely to be double-counted demand. Watch for leakage: if the held-out group can still be reached by other campaigns, the baseline is contaminated.
Full definitionMulti-Touch Attribution
Assigns fractional conversion credit across every tracked touchpoint in an individual user's path, using observed journeys rather than aggregate modeling.
It stitches user-level events across channels and devices, then applies a rule-based or algorithmic model to divide credit along each path. Advertisers want it because it operates at campaign and keyword granularity, which MMM cannot. Its foundation has eroded: cross-site identifiers, app tracking consent, and walled-garden reporting all break the stitching, so paths are increasingly incomplete and credit gets concentrated on the channels that still report cleanly.
Full definitionSide by side.
The differences that actually change what happens in your account.
| Holdout Testing | Multi-Touch Attribution | |
|---|---|---|
| What it manipulates | Exposure - a slice of the audience is actually denied ads. | Nothing - it only observes touches on paths that already happened. |
| Has a real counterfactual | Yes - the excluded group's behavior is the baseline. | No - there's no group that didn't see anything to compare against. |
| Grain of the answer | Whole audience or campaign the holdout was built against. | Down to individual campaign, ad group, or keyword. |
| Update cadence | Runs as a defined test window, or stays on as a standing check. | Continuous - updates as new tracked events flow in. |
| What tracking loss does to it | Makes it harder to keep the same people reliably excluded over time. | Breaks paths outright, dropping touches and shifting credit toward whichever channel still reports cleanly. |
| Main failure mode | Leakage - the excluded group gets reached anyway through a channel you don't control. | Path incompleteness - missing touches concentrate credit on the wrong channels. |
| Best question it answers | Would this specific audience have converted without seeing ads at all. | Which touchpoints appeared most often on paths that converted. |
What actually separates them.
Holdout testing has a built-in counterfactual because part of the audience never saw ads; MTA has none, since every path it analyzes already converted or didn't, with no controlled non-exposure to compare against.
MTA operates continuously on whatever tracking flows in, while a holdout has to be deliberately built and, in most cases, deliberately maintained to keep the same people excluded.
Tracking loss degrades the two differently - it makes holdout enforcement harder (you can't keep verifying who's excluded), while it makes MTA's paths shorter and more biased toward last-touchable channels.
Holdout testing only ever answers for the one audience it was built on, while MTA reports across every campaign and channel simultaneously, just without any causal grounding behind the numbers.
A holdout can return a genuinely flat result - no difference between exposed and excluded - which is informative; MTA structurally cannot return a flat result, because it always distributes 100 percent of credit somewhere even when the real incremental effect is near zero.
Which one should you use?
Use Holdout Testing when
- You want to know whether a specific audience, like retargeting or CRM, would have converted without seeing the ads.
- You have a large enough named list and the ability to reliably enforce suppression on it.
- You're willing to wait for a defined test window, or want a standing baseline that runs indefinitely.
- You need a defensible causal answer, not a directional one, for a spend decision on that audience.
Use Multi-Touch Attribution when
- You need to compare campaigns, ad groups, or keywords against each other for a routine, near-term budget shift.
- You want continuous, up-to-date reporting rather than waiting on a test window to close.
- Your tracking is intact enough that paths aren't badly broken by cookie loss or app tracking consent.
- You're using the output directionally, alongside a periodic holdout, rather than as your only measurement source.
Common questions.
If MTA already shows my retargeting campaign converting well, do I still need a holdout?
Yes, because MTA is just describing who touched what before converting, not whether the ad caused the conversion. Retargeting is exactly the case where MTA tends to overstate value, since it often reaches people who were already close to buying, and only a holdout can tell you if that's happening.
Why does my holdout audience keep shrinking or losing integrity over time?
This is usually leakage - other campaigns, channels, or even organic search are still reaching people you meant to exclude, or the identifier you're using to enforce the holdout isn't stable across devices and sessions. Both erode the baseline, so it's worth periodically checking that the excluded group is actually staying excluded.
Can I use MTA data to decide which audience to holdout-test next?
Yes, that's a reasonable use - if MTA shows a channel or audience getting a disproportionate share of credit relative to its actual media weight, that's a good candidate to verify with a holdout, since MTA's credit allocation is exactly the kind of number a holdout can confirm or debunk.
Or stop choosing between them.
AdFlint picks the setting, writes the ads, and keeps optimizing inside the Google and Meta accounts you already own.
Related comparisons
- A/B Testing vs Incrementality Testing
- A/B Testing vs Conversion Lift Tests
- A/B Testing vs Brand Lift Tests
- A/B Testing vs Geo Testing
- A/B Testing vs Holdout Testing
- A/B Testing vs Marketing Mix Modeling
- A/B Testing vs Multi-Touch Attribution
- Conversion Lift Tests vs Incrementality Testing