01
Decide what you are asking, before you touch anything
A geo holdout answers one question: if we stop spending here, how much revenue do we lose? Not which creative is better, not which audience converts, not whether the channel is cheap. One channel, one question, one number. Write the question down and the size of result that would change your decision. If a 10% lift would not change what you do, the test is not worth running.
02
Match the markets before you split them
Pick pairs of markets that already behave alike. Take at least eight weeks of history and pair regions whose weekly revenue tracks each other closely, then assign one of each pair to holdout and one to control. Matching on population or on store count is the common mistake: what has to match is the shape of the revenue line, because that is what you will be comparing.
03
Check the pre-period actually tracks
Before the test starts, run an equal-length pre-period with nothing changed and confirm the two groups move together within a few percent. If they diverge while you are doing nothing, they will diverge during the test too, and you will read that drift as an advertising effect. A failed pre-period is a cheap way to find out the test would have lied to you.
04
Work out the smallest effect you could detect
This is the step almost everyone skips, and it is the one that decides whether the test can work at all. Given your revenue volume, your week-to-week variance and the number of market pairs, there is a minimum effect size below which the noise swallows the signal. Compute it first. If the smallest detectable effect is 30% and you expect 10%, the test will come back inconclusive no matter how carefully you run it, and you will have spent the money to learn nothing.
05
Turn it off cleanly, and leave it off
Full suppression in the holdout markets, for the whole run. Two to four weeks for a high-volume channel with eight or more pairs; four to six for lower volume. Extend the duration rather than cutting the number of pairs, because pairs buy you precision and days only buy you sample. And do not change anything else: no new creative, no promotion in one region, no budget shift elsewhere. One change at a time is the whole method.
06
Read the interval, not the headline
The result is a gap between holdout and control, and it comes with a confidence interval. A lift of 35% with an interval of 22% to 49% is a real effect, and the honest way to report it is as the range. A lift of 5% with an interval of minus 8% to plus 18% includes zero, which means the test did not settle it. Reporting that second one as 5% incremental is the single most common way a good test becomes a bad decision.
Why this beats the dashboard
Platform reporting cannot tell you what would have happened anyway. It can only tell you which sales it can plausibly claim, which is why the platforms together claim more sales than you made. A holdout measures the counterfactual directly. It is slower and more expensive, and it is the only number in marketing measurement that survives being argued with.
Decifer runs this method in the Analytics and Paid modules, including the minimum detectable effect check before a test is worth starting.