Testing & Measurement

Holdout Testing vs Marketing Mix Modeling: One Audience or the Whole Mix

In short: Holdout testing is a live, randomized experiment on one specific audience; Marketing Mix Modeling is a backward-looking regression across your entire channel mix. Holdout testing suppresses ads to a slice of a named list - usually retargeting or CRM - and reads the gap in behavior against everyone else as the incremental effect. MMM never touches delivery; it fits historical sales against spend, seasonality, and price across every channel at once, including offline media a holdout could never reach. Holdout testing gives you a trustworthy causal number for one audience; MMM gives you a simultaneous, if correlational, estimate for the whole mix. Use holdout testing to police a specific suspect audience, and MMM to set the overall budget split across channels you can't run individual experiments on.

By the AdFlint research team · Fact-checked against current Google and Meta platform behavior · Last reviewed July 2026

Holdout Testing

Permanently excludes a randomly selected share of your audience from seeing ads, so their behavior serves as an ongoing baseline.

A fixed percentage of users or accounts is suppressed from targeting for the duration of the test, and the gap between held-out and exposed groups gives you the incremental effect. Advertisers keep a standing holdout on retargeting and CRM audiences, where reported returns are most likely to be double-counted demand. Watch for leakage: if the held-out group can still be reached by other campaigns, the baseline is contaminated.

Full definition

Marketing Mix Modeling

Estimates each channel's contribution by fitting aggregate sales against spend, seasonality, pricing, and external factors over time, without user-level data.

A regression over historical weekly or monthly data separates the effect of each media channel from seasonality, promotions, and macro noise, and can express diminishing returns and carryover. It appeals because it needs no cookies or identifiers and covers offline media. It demands years of clean history, cannot guide day-to-day bidding, and is correlational, so standard practice is to calibrate the model against real incrementality experiments.

Full definition

Side by side.

The differences that actually change what happens in your account.

 Holdout TestingMarketing Mix Modeling
What's being testedOne named audience - retargeting pool, CRM list, lookalike segment - split into exposed and excluded groups.Every channel in the media mix simultaneously, using historical spend and sales.
Requires changing deliveryYes - a portion of the audience is actually blocked from seeing ads.No - it's fit entirely on data from campaigns that already ran normally.
Identity requirementNeeds a stable identifier to keep the same people excluded for the whole run.None - works on aggregate spend and sales totals, no individual identifiers at all.
Offline channel coverageNone - it only works on channels where you can target and suppress specific people.Full - covers TV, print, and other offline media as long as spend history exists.
Causal or correlationalCausal - the excluded group is a real counterfactual for that audience.Correlational - infers contribution from patterns and needs calibration to be trusted.
Typical use caseChecking whether retargeting or CRM spend is buying incremental buyers or just people who'd convert anyway.Setting the overall budget allocation across the full channel mix for a coming period.
Time horizonCan run as a standing, ongoing check indefinitely.Needs years of history to fit and is typically refreshed quarterly.
Failure modeLeakage - the excluded group still gets reached through a channel you don't control.Multicollinearity - channels that move together can't be cleanly separated.

What actually separates them.

01

Holdout testing produces a real counterfactual by physically withholding ads from part of one audience; MMM never withholds anything, it only observes what already happened.

02

Holdout testing can only ever speak to the one audience it was built on, while a single MMM run produces an estimate for every channel in the account at once.

03

MMM can incorporate channels with no addressable audience at all, like linear TV or print, which a holdout can never touch since it requires suppressing specific identifiable people.

04

A holdout's result is a trustworthy number for a narrow question; feeding that result into MMM as a calibration point is standard practice for keeping MMM's coefficients honest on the channels a holdout can reach.

05

Holdout testing needs an identifier to keep the same people excluded over time, while MMM needs no identifiers at all, only spend and sales totals - so MMM survives even for channels holdout testing structurally can't measure.

Which one should you use?

Use Holdout Testing when

  • You suspect a specific audience - retargeting, CRM, an email list - is mostly converting on its own regardless of the ads.
  • You have a large enough named list to split with statistical power and a way to reliably keep the excluded group excluded.
  • You want an ongoing, standing check rather than a one-time strategic study.
  • The channel in question is addressable - you can actually target and suppress specific identified people.

Use Marketing Mix Modeling when

  • You need one number for how your total budget should split across every channel, including offline media.
  • You have years of clean spend and sales history and want to model diminishing returns across the mix.
  • Some of your channels - TV, radio, sponsorships - have no addressable audience a holdout could ever be built on.
  • You're setting quarterly or annual budget, not checking one audience's incrementality this month.
  • You already have holdout or geo test results on a few channels and want to use them to keep the model's coefficients grounded.

Common questions.

Can I use my holdout test results to improve my MMM?

Yes, and it's considered good practice. Feeding a holdout's causal incrementality number in as a calibration point for the channel it covers helps anchor MMM's coefficient for that channel instead of leaving it purely inferred from correlation.

Why can't I just run a holdout on my TV spend?

Holdout testing requires suppressing ads to specific identified people, and TV bought at a market or network level doesn't let you target or exclude individuals that precisely. That's exactly the gap MMM is built to fill, since it works on aggregate spend without needing to address anyone individually.

My holdout showed almost no lift on retargeting - does that mean I should cut it entirely?

It means that audience wasn't buying much incrementality, which is common on retargeting since it often reaches people already close to converting. Whether to cut it depends on the actual gap size versus the spend, but a near-zero read is a strong signal to reduce spend on that audience rather than ignore the result.

How often should I refresh my MMM if I'm also running holdouts?

Most shops refresh MMM quarterly regardless, since that matches typical budget planning cycles, but a standing holdout can feed fresh calibration data into every refresh rather than waiting for a one-off study. That keeps the model from drifting too far from what a live experiment is actually showing.

Or stop choosing between them.

AdFlint picks the setting, writes the ads, and keeps optimizing inside the Google and Meta accounts you already own.

Related comparisons

All comparisons