A/B Testing vs Holdout Testing
In short: Both split an audience, but one is a short comparison and the other is a standing baseline. A/B testing runs for a defined window to pick a winner between two treatments, then ends. Holdout testing permanently excludes a slice of the audience from targeting so their behavior keeps serving as an ongoing baseline for incremental impact. A/B testing answers which version wins; holdout testing answers whether this audience segment is worth targeting at all, continuously. If you need a one-time decision between options, A/B test; if you need a standing check on whether a segment like retargeting or CRM is adding real conversions, run a holdout.
By the AdFlint research team · Fact-checked against current Google and Meta platform behavior · Last reviewed July 2026
A/B Testing
Splits a randomized audience between two variants that differ in one element, then compares outcomes to decide which performs better.
Both platforms provide built-in split tests that divide the audience so people see only one cell, avoiding the overlap you get from simply running two ad sets side by side. Use it for creative, landing pages, audiences, or bid strategy, one variable at a time. The usual failure is calling a winner on a handful of conversions, where the observed gap is well inside normal variance.
Full definitionHoldout Testing
Permanently excludes a randomly selected share of your audience from seeing ads, so their behavior serves as an ongoing baseline.
A fixed percentage of users or accounts is suppressed from targeting for the duration of the test, and the gap between held-out and exposed groups gives you the incremental effect. Advertisers keep a standing holdout on retargeting and CRM audiences, where reported returns are most likely to be double-counted demand. Watch for leakage: if the held-out group can still be reached by other campaigns, the baseline is contaminated.
Full definitionSide by side.
The differences that actually change what happens in your account.
| A/B Testing | Holdout Testing | |
|---|---|---|
| Duration | Fixed window, typically days to a couple weeks, then it ends and you act on the winner. | Ongoing - the excluded slice stays suppressed indefinitely as a standing baseline. |
| What is excluded | Nothing - every user in the test sees some ad, just a different variant. | A fixed percentage of users or accounts is suppressed from targeting entirely for as long as the test runs. |
| Typical use case | Picking between two creatives, landing pages, or bid settings. | Checking whether a segment like retargeting or CRM audiences is generating conversions beyond what would happen anyway. |
| What you get at the end | A declared winner you roll out to all traffic. | A continuously updated gap between held-out and exposed groups, re-readable at any time. |
| Leakage risk | Audience overlap between the two test cells. | The held-out group still gets reached through another campaign or channel, contaminating the baseline. |
| Resource cost | Low - runs inside normal campaign delivery and budget. | Higher - you are deliberately not marketing to a slice of valuable audience for as long as the holdout runs. |
| Where it reports | Platform's built-in experiment dashboard, resolved once significance is reached. | A recurring comparison pulled on demand, since there is no defined end date. |
What actually separates them.
A/B testing has a defined end date and produces a one-time winner; holdout testing has no end date by design and produces a continuously refreshable baseline.
In A/B testing every cell is exposed to some ad; in holdout testing one cell is deliberately excluded from all targeting for the segment being tested.
A/B testing is cheap to run constantly because both cells still convert; holdout testing has an ongoing cost, since the held-out slice is valuable audience you are choosing not to market to.
Leakage in an A/B test means the two cells overlapped; leakage in a holdout means the excluded group got reached anyway through a different campaign, quietly erasing the baseline.
A/B testing is scoped to one campaign's variable; holdout testing is usually scoped to an audience segment, like retargeting or CRM lists, that spans multiple campaigns.
Which one should you use?
Use A/B Testing when
- You need to choose between two specific treatments and move on.
- The decision is scoped to one campaign, not a whole audience segment.
- You want a result within a couple of weeks, not an open-ended commitment.
- Both variants can run at full budget without one deliberately going dark.
Use Holdout Testing when
- You want a permanent, always-available baseline for a segment like retargeting or CRM audiences.
- Your reported returns on a segment look suspiciously good and you want an ongoing check, not a one-time read.
- You can absorb the cost of permanently not marketing to a slice of a valuable audience.
- You need to catch incrementality drift over time, not just a single snapshot.
A/B Test Significance Calculator
Enter visitors and conversions for two variants to get conversion rates, uplift, z-score, p-value, and whether the result is significant.
Open the free calculatorCommon questions.
How big should a holdout group be?
Big enough to detect a realistic lift without starving your working audience of budget, which for most advertisers lands the held-out slice at a modest, fixed percentage rather than an even split. The exact size trades off statistical confidence against how much valuable audience you are willing to leave untargeted.
Can I turn a holdout test into an A/B test?
Not directly, since a holdout compares exposure against no exposure while an A/B test compares one treatment against another, but you can run an A/B test on the exposed side of a holdout to optimize creative for the group you are still targeting. That gives you both an incrementality baseline and a performance comparison at the same time.
Why does my holdout group keep converting even though they see no ads?
That is the whole point - some conversions happen regardless of advertising, whether from brand recognition, organic search, or repeat purchase behavior, and the holdout's converted volume is your estimate of that baseline. The gap between the held-out and exposed groups, not the held-out group's raw conversion count, is what tells you the ad's incremental effect.
What breaks a holdout test?
Leakage is the main risk. If the held-out group can still be reached through another campaign, a different platform, or a broad audience that overlaps with the excluded segment, the baseline stops being clean and the measured lift understates the true effect.
Or stop choosing between them.
AdFlint picks the setting, writes the ads, and keeps optimizing inside the Google and Meta accounts you already own.
Related comparisons
- A/B Testing vs Incrementality Testing
- A/B Testing vs Conversion Lift Tests
- A/B Testing vs Brand Lift Tests
- A/B Testing vs Geo Testing
- A/B Testing vs Marketing Mix Modeling
- A/B Testing vs Multi-Touch Attribution
- Conversion Lift Tests vs Incrementality Testing
- Brand Lift Tests vs Incrementality Testing