Holdout Tests: What a Trustworthy Marketing Result Requires
A case study says sales rose sharply after the campaign launched. The number may be accurate and still not say what the campaign did, because nothing on the page shows what those sales would have done anyway. Here is the design built to supply that comparison, and what it must disclose to count as evidence.
The problem every before-and-after case study shares
A before-and-after claim compares one account, brand, or set of stores to its own past. No group in that comparison went without the campaign, so everything else that moved in the window gets folded in and reported as campaign effect. The season turns, a competitor leaves, the category grows, a weak quarter reverts toward normal, and each produces the shape a results page wants, a line climbing after the start date. This is the number no ordinary case study can show. The gap is structural, not dishonest: a group measured against itself cannot separate the campaign from the calendar.
What “holding a group back” actually means
The fix is a comparison group that never receives the campaign. Experimental design calls this a control group: units receiving a standard treatment, a placebo, or no treatment at all.
A holdout test excludes a comparable slice of the same audience. That slice sees nothing, the treated slice sees everything, and both are measured on the same metric across the identical window. A matched-market test moves this into geography: the campaign runs in one city or set of stores and is withheld from a comparable one over the same period. A Google Research paper on geo experiments describes them as non-overlapping regions randomly assigned to a control or treatment condition.
The identical window is not optional. Compare a treated group’s third quarter against a holdout’s second and you reintroduce the seasonality the design removed.
Why “comparable” is doing most of the work
A control group only works as a counterfactual if it resembles the treated group before the campaign starts. A holdout that began smaller, skewed to a different customer type, or carried different prior purchase behaviour measures the difference between two populations, with a campaign on top.
So the assignment method belongs in the write-up. Random assignment is the strongest version, because randomized experiments allow the greatest reliability in estimating treatment effects: nobody chose which group got the campaign, so nothing systematic hides in the split. A weaker but honest version names the matching variables and how markets were paired. The weakest version that still calls itself a test is “we selected two similar markets.” When assignment goes undisclosed, treat the comparability as unverified rather than the result as false, the standard applied to what backs up a case study number anywhere else.
A hypothetical worked example: two towns, one campaign
Everything below is hypothetical. It describes no real agency, brand, or campaign, and the figures only show the mechanics.
Picture a bakery chain with outlets in two towns of similar size, each selling about 100 units in a normal week. The chain runs a campaign in Town A for eight weeks and nothing in Town B across the same eight weeks. Town A ends at 130 a week. A before-and-after write-up stops there and reports a 30 percent rise. But Town B ends at 120. Something lifted both towns. The campaign’s estimated contribution is the gap between the two changes, roughly 10 percentage points, not the 30 the treated town shows alone.
Those figures are illustrative arithmetic and nothing else. No invented example should ever be phrased so it could pass for a reported result.
What can still go wrong with a well-designed test
- Contamination. The holdout sees the campaign anyway, through another channel or a retargeting list nobody split. Experimental design calls this diffusion: treatment effects spread into the control group, and the measured difference shrinks even though the campaign worked.
- Spillover between markets. A spillover is an indirect effect on a subject the experiment never treated. Customers commute, order online across boundaries, and see organic content that ignores every line the media plan drew.
- A window too short, or a sample too small. If an effect takes ten weeks to reach purchase behaviour and the test ran four, it reports no lift for a campaign that had one. In small groups, ordinary weekly wobble outweighs the gap.
- Publishing only the winning pair. Run the design across six market pairs and one looks best by chance. Shelving the other five is invisible unless the write-up says how many ran.
A short checklist for a claimed holdout or matched-market result
- Does it say what defines the comparison group and how members or markets were assigned?
- Were both groups measured across the exact same calendar dates? Look for two date ranges stated, not implied.
- Is the compared metric the same on both sides? Sales on one side and site visits on the other is not a comparison.
- Does it disclose the size of both groups, not just the treated one? A tiny holdout gives a noisy baseline.
Why this standard shows up so rarely
The honest reason is cost, not evasion. A holdout means withholding a campaign someone believes in from part of a paying client’s audience. A matched-market test needs two comparable markets, coordinated buying, and clean sales data. What it yields is a modest gap between two lines, not a soaring arrow. So the absence of a holdout does not make a case study worthless. It makes the number suggestive rather than conclusive, and that distinction is the takeaway, not a verdict on any particular agency.
Read the library with this in mind
Browse the case study library and check what each write-up says about its comparison group, if it names one. If you have published work of your own, submit your own with the assignment method and both group sizes stated.
FAQ
What is a holdout test in marketing?
A holdout test withholds a campaign from a comparable slice of the audience, so that slice’s outcome over the same window stands in for what would have happened without it. What matters is the difference between the groups’ changes, not the treated group’s alone.
Is an A/B test the same as a holdout test?
They share the logic of a comparison group but answer different questions. An A/B test compares two live versions of something to the same audience segment, asking which execution wins. A holdout compares campaign against nothing. Our piece on reading an A/B test claim covers that adjacent case, and how to read a marketing case study without being misled sets out the wider framework.
Sources
- Treatment and control groups, Wikipedia, fetched 7 September 2026.
- Randomized experiment, Wikipedia, fetched 7 September 2026.
- Measuring Ad Effectiveness Using Geo Experiments, Google Research, fetched 7 September 2026.
- Internal validity, Wikipedia, fetched 7 September 2026.
- Spillover (experiment), Wikipedia, fetched 7 September 2026.