Geo Split vs Full Pause: Designing a Brand-Paid Holdout That Survives Finance Review

The two viable test designs for measuring branded-paid incrementality — a geographic holdout versus a full national pause — compared on statistical power, contamination risk, and how each survives finance review.
Geo Split vs Full Pause (Brand-Paid Holdout Design)
Two test designs for measuring branded paid search incrementality: pause brand in matched regions (geo split) or pause nationally on a time window (full pause).
A brand-paid holdout test answers one question: how much of the revenue attributed to your branded Google Ads campaigns would have arrived anyway through organic? Two designs are defensible. A geo split pauses brand keywords in a set of treatment DMAs while keeping them live in a matched control set, then compares revenue trajectories. A full pause turns brand off nationally for a defined window and compares against a pre-period baseline. Geo splits give you a concurrent counterfactual and lower revenue risk; full pauses give you higher statistical power but leave seasonality as the elephant in the finance meeting.
The choice between the two isn't really methodological — it's political. A geo split is what your measurement lead wants because it isolates the treatment effect against a concurrent control. A full pause is what your CFO reluctantly signs off on because it's simpler to explain and reads cleanly on a P&L timeline. Both can produce a credible incrementality number. Only one of them will survive when a competitor decides to bid on your brand term the week you're in-market.
Before you pick, three upstream questions determine which design is even viable: does your organic brand listing sit above the fold on the relevant SERPs, is any competitor currently conquesting on your brand term, and can you tolerate 2-4% of quarterly revenue potentially disappearing during a national pause? If any of those fail, the design choice is made for you.
Geo split vs full national pause — design trade-offs at a glance
| Dimension | Geo split holdout | Full national pause |
|---|---|---|
| Counterfactual | Concurrent control DMAs | Pre-period baseline |
| Statistical power (typical MDE) | 6-12% lift detectable | 3-6% lift detectable |
| Duration to read | 21-35 days | 10-21 days |
| Revenue at risk | 0.5-1.5% of quarterly revenue | 2-4% of quarterly revenue |
| Seasonality confound | Controlled by design | Requires modelled baseline |
| Contamination risk | High (spillover, national campaigns) | Low (nothing to spill into) |
| Setup complexity | Medium-high (DMA matching, geo targeting) | Low (single global pause) |
| Finance sign-off difficulty | Medium | High |
| Repeatability per year | 3-4 tests | 1-2 tests |
Notice the tension in that table. Full pauses win on power and speed but lose on risk and seasonality defense. Geo splits win on risk and concurrent control but require more setup and more time in-market to reach significance. The right design depends on which trade-off your organisation can actually absorb — not which one the measurement blog post prefers.
Geo split: when the concurrent counterfactual is worth the setup cost
A geo split test pauses branded paid search in a treatment set of DMAs (typically 8-15 markets covering 20-30% of national revenue) while leaving the campaigns live in a matched control set. Because both sets run concurrently, any macro shock — a competitor promo, an economic wobble, a viral moment — hits both groups. What differs between them is attributable to the pause.
The setup work matters. Matching treatment and control DMAs on pre-period revenue trend, brand search volume, and customer mix is the single biggest determinant of whether the test reads cleanly. A synthetic-control approach or a straightforward pre-period regression on 12-16 weeks of daily data usually gets you there. Skip this and your minimum detectable lift balloons past the point where the test can conclude anything.
Watch for organic spillover contamination
Geo splits leak in ways full pauses don't. National display campaigns, TV bursts, email sends, and even organic social don't respect DMA boundaries — so a treatment DMA still gets exposed to brand demand generation. If your national brand marketing spend is heavy during the test window, the incrementality number you read is a floor, not a point estimate. Plan the holdout for a period of stable national activity, or account for spillover explicitly in the read.
Full national pause: higher power, higher stakes, longer arguments with finance
A full national pause is simpler and statistically stronger. You turn brand paid off everywhere for a defined window — typically 14-21 days after accounting for ramp and decay — and compare revenue against a modelled baseline drawn from the prior 8-12 weeks and the same window in prior years. Because 100% of traffic is in the treatment, you detect smaller lifts faster.
The problem is defensibility. When you present a 4% incrementality read to the CFO, the first question is 'how do we know it wasn't seasonality?' Your answer has to be a modelled counterfactual — same weeks prior year, weather-adjusted, promo-calendar-adjusted — and it has to be built before the test starts. Retrofitting a seasonality defense after the fact is how good measurement work gets dismissed as post-hoc rationalisation. Sizing revenue-at-risk explicitly before you start is also non-negotiable; a national pause with no floor plan is how measurement teams lose executive trust.
Statistical power vs revenue-at-risk by test design
Minimum detectable lift (%)
Revenue at risk (% of quarterly revenue)
Frequently asked questions
Start with a geo split. The revenue-at-risk is bounded, the concurrent control gives you a defensible counterfactual, and if the read is inconclusive you can escalate to a full national pause with a much stronger internal case. Running a national pause as your first-ever brand incrementality test is how measurement programs get shut down after one bad month.
For a geo split, seasonality is controlled by design — the control DMAs experience the same seasonality as the treatment set. For a full national pause, you need a pre-registered baseline model built from at least 8-12 weeks of pre-period data plus year-over-year same-window comparison, ideally adjusted for known promo calendars. Building that model after seeing the results is not defensible to a serious finance team.
Typically 8-15 treatment DMAs matched to an equal-sized control set, together covering 20-30% of national revenue. Fewer than 6 DMAs per arm and the minimum detectable lift usually exceeds 15%, which is larger than most branded-paid incrementality effects — the test will read null even when the effect is real.
Discard the first 3-5 days as ramp (users who saw the ad in the prior week are still converting) and the last 2-3 days as decay if you're re-enabling. That means a 21-day pause gives you roughly 13-15 clean measurement days. Full national pauses can often read at 14 days total; geo splits usually need 21-28.
This is the failure mode that invalidates the most branded-paid holdouts. Monitor branded impression share and top-of-page rate daily during the holdout — if a competitor enters, the test measures a different question (brand paid vs competitor-conquested organic) and the incrementality read is no longer generalisable. Have a stop-test rule pre-registered before you start.
For most stores, brand paid represents 8-15% of paid revenue and 15-30% of that is genuinely incremental, so a 14-day pause typically costs 0.4-1.2% of quarterly revenue in genuinely lost sales. That's the number to size before the test, not the gross revenue running through brand paid — which conflates cannibalised organic with true incrementality.
You can, and it lowers revenue-at-risk proportionally to Bing's share of your branded search. But the read doesn't transfer cleanly — Bing branded search is a different user population (older, more desktop-heavy), so the incrementality percentage you measure there is not the number you should apply to Google budgets. Use it as a directional pilot, not a final answer.
Geo splits can be re-run 3-4 times a year by rotating treatment and control DMA assignments; this also helps average out any residual match imperfection. Full national pauses are usually a once- or twice-yearly event because of the revenue disruption and the finance conversation each one requires.
Typical DTC results land in the 10-35% incremental range — meaning 65-90% of branded-paid revenue would have arrived through organic anyway. The exact number depends heavily on SERP layout: if your organic listing sits below shopping tiles and a Wikipedia panel, incrementality skews toward the top of that range. If it sits at position one above the fold, it skews toward the bottom.
The same design logic applies but the mechanics differ. Meta doesn't offer DMA-level targeting with the same granularity, so geo splits are harder to instrument cleanly. A structured pause of branded audiences on Meta is more commonly run as a matched-market test using platform-native lift studies rather than as a self-instrumented holdout.
Track CAC, channels, and funnel conversion in one place
Metricuno connects ad spend, funnel events, and revenue so you can see CAC by channel, cohort, and campaign — without stitching together five tools.