Geo Experiments: The Complete Guide to Measuring Marketing Incrementality
Your ad platform says the campaign delivered a 4.2 ROAS. Your finance team looks at total revenue and sees no change. Both of them are reading real numbers. Only one of them is measuring what the campaign actually caused. This gap is the single most expensive measurement problem in marketing, and it exists because attribution answers the wrong question. Attribution asks which touchpoint preceded the conversion. The question you actually care about is what would have happened if we hadn't run this at all. Those are different questions, and no amount of clean tracking will turn the first into the second. Geo experiments are the most practical way to answer the second question at scale. This guide covers what they are, when they earn their keep, how to design and analyze one properly, and where they fall short.

Your ad platform says the campaign delivered a 4.2 ROAS. Your finance team looks at total revenue and sees no change. Both of them are reading real numbers. Only one of them is measuring what the campaign actually caused.
This gap is the single most expensive measurement problem in marketing, and it exists because attribution answers the wrong question. Attribution asks which touchpoint preceded the conversion. The question you actually care about is what would have happened if we hadn't run this at all. Those are different questions, and no amount of clean tracking will turn the first into the second.
Geo experiments are the most practical way to answer the second question at scale. This guide covers what they are, when they earn their keep, how to design and analyze one properly, and where they fall short.
What is a geo experiment?
A geo experiment is a randomized or quasi-randomized controlled trial where the unit of assignment is a geographic region rather than an individual user.
You split your markets into two groups. One group (treatment) gets a change in marketing activity: a new channel switched on, budget increased, a campaign paused, a creative rotated in. The other group (control) carries on exactly as before. You then compare what happened to your business KPI in the treatment markets against a credible estimate of what would have happened there without the change.
That estimate of the unobserved alternative is the whole game. It is called the counterfactual, and constructing it well is what separates a geo experiment from a chart with two lines on it.
A few terms worth pinning down:
Term
What it means
Geo unit
The geographic building block you assign: DMA, state, province, city, ZIP/postal cluster
Treatment geos
Markets that receive the change in marketing activity
Control geos
Markets held at business as usual, used to build the counterfactual
Pre-period
Historical window used to establish that treatment and control move together
Lift
The difference between observed KPI in treatment and the estimated counterfactual
Incrementality
The share of observed outcomes that the activity actually caused
iROAS / iCPA
Incremental revenue per unit of spend, or incremental cost per acquisition
The critical property is that geo experiments measure at the market level, not the user level. You never need to know which individual saw the ad. You compare aggregate outcomes in two sets of regions. That one design choice is what makes geo experiments resilient to almost everything that has broken user-level measurement over the past five years.
Why geo experiments exist
Attribution is a correlation engine
Last-click, multi-touch, view-through: all of these describe the observed path to conversion. None of them can tell you whether the conversion would have occurred anyway. Users who search for your brand and click your ad were often already on their way to you. Attribution happily credits the ad. A geo experiment does not, because the control markets contain those same already-converting users.
The ad platforms grade their own homework
Platform-reported conversions are measured by the party selling you the media, using their own attribution windows and their own definition of a view. This is not necessarily dishonest, but it is structurally optimistic, and it is not comparable across platforms. A geo experiment produces one number measured in your own data, on your own KPI definition, in the same units for every channel.
Privacy changes removed the plumbing
ATT, cookie deprecation, consent modes, and modeled conversions have all made user-level tracking partial at best. Geo experiments are largely immune to this because they never depended on user-level identity in the first place.
Some channels were never trackable
CTV, podcast, radio, out-of-home, print, direct mail, sponsorships, influencer campaigns. If your measurement approach requires a click, these channels either look worthless or are invisible entirely. Geo experiments measure them exactly the same way they measure paid search.
Spillovers are captured, not lost
A user sees your billboard, tells a colleague, the colleague searches for you and converts organically. User-level tracking loses that conversion completely. Market-level measurement captures it, because it lands inside the treated region.
MMM needs external calibration
Marketing mix models are fitted to observational data, which means their channel effects can be poorly identified when channels move together. Geo experiment results provide causal anchors that can be used to calibrate an MMM, most explicitly in Bayesian frameworks where a measured effect becomes an informative prior. Google's Meridian is built around this pattern, and Google has announced Meridian GeoX, an open-source geo experiment solution designed to turn test results directly into MMM priors.
Experiments and MMM are not competing approaches. MMM gives you breadth across all channels continuously. Geo experiments give you depth and causal credibility on one question at a time. Used together, each fixes the other's main weakness.
When to use geo experiments
Reach for a geo experiment when most of these are true:
You can target geographically. The channel must let you turn activity up, down, or off by region. Most major platforms do. Some do not, or only at coarse levels.
You have enough geos. As a working minimum, roughly 20 to 30 assignable markets with meaningful volume, ideally more. With 210 US DMAs you have plenty of design freedom. With five European countries you do not have an experiment, you have an anecdote.
Your KPI is measurable by region and reasonably stable. Daily or weekly revenue, orders, installs, signups, or store visits, mapped to geography, with at least a year of clean history.
The decision is worth the cost of learning. A geo experiment usually means deliberately spending differently from the optimum in some markets for four to eight weeks. That opportunity cost only makes sense for budget decisions of real size.
The effect could plausibly be large enough to detect. Geo experiments detect market-level shifts. If a channel accounts for 0.5% of your revenue, you are unlikely to detect it above the noise no matter how well you design.
Concrete situations where geo experiments are the right tool:
- Validating whether an expensive channel is genuinely incremental, especially branded search, retargeting, and app install campaigns
- Measuring untrackable media: CTV, audio, OOH, sponsorships, direct mail
- Sizing the incremental value of a budget increase before rolling it out nationally
- Testing a pricing, promotion, or product change that cannot be randomized at user level
- Deciding whether to keep or kill a channel that attribution and MMM disagree about
- Generating calibration priors for an MMM
- Measuring a channel whose own reported numbers you have reason to distrust
When not to run one
Be equally clear about the cases where a geo experiment is the wrong instrument:
- You can randomize users instead. If the change is a website or in-product experience, run an A/B test. It will be faster, cheaper, and far more precise.
- Too few geos, or one market dominates. If a single city is 40% of revenue, assignment is unbalanced by construction and no clever weighting fully rescues it.
- The expected effect is small. Run the power analysis first. If the minimum detectable effect comes back at 30% and you expect 5%, you have learned something valuable at zero cost: don't run the test.
- You need an answer in ten days. Geo experiments need a pre-period, a treatment period, and often a cool-down. Four to eight weeks is typical.
- Your geo-level data is unreliable. Missing region tags, inconsistent attribution of orders to markets, or heavy VPN traffic will quietly break everything downstream.
- The market is in an unusual state. Launching mid-Black Friday, during a competitor's national campaign, or across a major regulatory change contaminates the comparison.
Geo experiments vs. A/B tests vs. MMM
A/B test
Geo experiment
MMM
Unit
User / session
Region
Aggregate time series
Causal strength
Highest
High
Correlational, model-dependent
Handles offline media
No
Yes
Yes
Privacy-resilient
Partially
Yes
Yes
Captures spillovers
No
Yes
Yes
Precision
High
Moderate
Low to moderate
Channel coverage
One at a time
One or a few at a time
All channels at once
Typical duration
Days to weeks
4–8 weeks
Continuous, refreshed quarterly
Main risk
Not applicable to media
Underpowered design
Identification and collinearity
The short version: A/B test what you can randomize by user, geo test what you can only randomize by market, and use MMM to hold the whole picture together.
The main geo experiment designs
Holdout (blackout) test. Turn the activity off in a subset of markets, leave it on elsewhere. This is the cleanest measure of incrementality because it directly estimates what you lose without the channel. It is also the most uncomfortable, since you are knowingly forgoing revenue in the treated markets. Best for the question "is this channel doing anything at all?"
Scale-up (lift) test. Increase spend in treatment markets while holding control flat. Lower business risk, and it measures the marginal return on additional budget rather than average return on all of it. Best for the question "should we spend more here?" Be aware that the change must be big enough to move the KPI. A 10% budget bump usually is not.
Matched market test. Pair markets that behave similarly in the pre-period and randomize within pairs. Improves precision considerably when the number of geos is limited.
Multi-cell design. Several treatment groups against a common control, for example three different spend levels. More expensive per cell, but you recover a response curve rather than a single point estimate, which is far more useful for budget planning.
Switchback / crossover. Alternate treatment and control status over time within the same markets. Efficient when you have few geos, but only valid when carryover effects are short relative to the switching interval, which for brand media they usually are not.
How to run a geo experiment, step by step
1. Write down the decision rule before you run the test
Not "does CTV work?" but "if incremental CPA on CTV is below $60, we shift $2M into it next year; above $90, we cut it." Pre-committing to the decision rule prevents the single most common failure mode, which is reinterpreting an ambiguous result to match what the team already wanted to do.
2. Choose the KPI, and choose it in your own data
Pick one primary business metric measured in your system: revenue, orders, first purchases, qualified signups. Not platform-reported conversions. Define one primary metric and treat everything else as secondary and directional.
3. Assemble geo-level history
You want at least 12 months of daily or weekly KPI data by geo unit, plus spend by geo and channel. Check for missing regions, sudden definition changes, and markets with erratic reporting. This step routinely takes longer than the test itself, and it is where most projects quietly go wrong.
4. Pick the geo unit
DMAs are the standard in the US because media buying aligns with them. Elsewhere, use states, provinces, or metro clusters. Smaller units mean more of them, and more units mean better statistical power, but only up to the point where media targeting can actually respect the boundary and where cross-border contamination stays modest.
5. Run a power analysis. Do not skip this.
This is the step teams skip, and it is the reason so many geo experiments end in a shrug. Before assigning anything, use your historical data to answer: given this design, this duration, and this budget change, what is the smallest true effect we could reliably detect?
Good practice is simulation-based. Take the pre-period data, inject synthetic effects of known size, run your intended analysis, and see how often you recover them. That gives you an honest minimum detectable effect. If the MDE exceeds what you plausibly believe the effect to be, redesign: more geos, longer duration, bigger spend delta, or a different question. Killing an underpowered test at the design stage is the highest-ROI thing a measurement team does all year.
6. Assign treatment and control
Randomize where you can, ideally within matched pairs or strata built on pre-period KPI level and trend. Where the number of geos is small, algorithmic market selection (as implemented in GeoLift and similar tools) searches for the treatment set whose synthetic control fits best historically.
Whatever the method, validate the same way every time: the treatment group and its counterfactual must track each other closely across the entire pre-period, and the gap between them must be stable rather than drifting. Also check that the two groups are comparable on the things that drive your business, such as seasonality shape, device mix, and channel composition.
7. Set an intervention large enough to detect
A frequent disappointment: a test is run with a spend change too small to produce a measurable KPI change, and the null result gets misread as "the channel doesn't work." Size the intervention off the power analysis, not off what feels comfortable.
Guardrails to set at the same time:
- Lock all other campaign changes for the test window, in both groups
- Watch for budget leakage, where paused markets push spend into control markets and contaminate them
- Beware platform auto-optimization, which may reallocate delivery across geos on its own
- Consider buffer zones between treatment and control markets that share media or commuting patterns
8. Run it, and leave it alone
Pre-register the design: geos, dates, KPI, analysis method, decision rule. Then do not peek and stop early. Sequential looks at a fixed-horizon test inflate false positives badly, and geo experiments are especially vulnerable because early days are noisy and the temptation to react is high.
Plan for a cool-down window after the treatment period. Brand and upper-funnel media have carryover: effects persist after spend stops, and they also ramp slowly after it starts. Ending measurement on the last day of spend usually understates the total effect.
9. Analyze with an explicit counterfactual
The standard toolkit, in rough order of sophistication:
- Difference-in-differences. Simple, transparent, and adequate when groups have genuinely parallel pre-trends. Fragile when they don't.
- Synthetic control. Builds a weighted combination of control geos that reproduces the treatment group's pre-period behavior, then projects it forward. This is the workhorse method and the basis of Meta's open-source GeoLift.
- Bayesian structural time series. Google's CausalImpact approach models the counterfactual as a time series with covariates and returns full credible intervals for cumulative effect.
- Time-based regression (TBR). Uses the pre-period relationship between treatment and control to predict the test window, well suited to designs with fewer geos.
Always report an interval, never a single number. And always run placebo checks: apply the identical analysis to a fake test date in the pre-period, and to control geos pretending to be treated. If you "find" effects where none can exist, your model is fitting noise and the real result cannot be trusted.
10. Convert to a business number, then decide
Translate lift into incremental ROAS or incremental CPA, compare it to your pre-registered threshold, and act. Then feed the estimate into your MMM as a calibration prior so a single 6-week test keeps paying off across every subsequent budget cycle.
A worked example: are app install campaigns incremental? (illustrative)
A subscription app tests whether its app install campaigns are incremental.
- 210 DMAs, 40 assigned to treatment via matched-pair randomization
- Design: holdout. Campaigns paused in treatment markets for 6 weeks, plus a 2-week cool-down
- Power analysis MDE: 8% change in new subscriptions
- Result: subscriptions in treatment markets came in 14% below the synthetic control, 95% interval 9% to 19%
- Spend saved during the holdout: $840K. Subscriptions lost: 6,200
- Implied incremental CPA: about $135, against a platform-reported CPA of $61
The decision was not to cut the channel. It was to stop planning against a $61 CPA. Budget was reallocated toward the segments where a follow-up multi-cell test showed marginal CPA still below the $110 payback threshold.
Pros of geo experiments
- Causal, not correlational. With a valid control group you are estimating an effect, not a coefficient in a model of your own choosing.
- Channel-agnostic. Offline, online, upper funnel, lower funnel, all measured on the same yardstick.
- Privacy-durable. No user identifiers, no cookies, no consent dependency.
- Captures the full effect, including word of mouth, cross-device journeys, halo onto organic and branded search, and offline conversions.
- Independent of the seller. Your data, your KPI, your definition of success.
- Calibrates everything downstream. One credible experiment can correct an entire MMM and reset internal ROAS benchmarks.
- Convincing to non-technical stakeholders. "We turned it off in 40 markets and here is what happened" travels much better in a board deck than a regression table.
Cons and limitations
- Statistically expensive. Your effective sample size is the number of geos, not the number of users. A test with 30 markets is a small-sample study however many millions of users sit inside those markets.
- Slow. Four to eight weeks per test, plus design and analysis time. You will run a handful per year, not hundreds.
- Real opportunity cost. Holdouts sacrifice revenue on purpose; scale-ups spend beyond the optimum on purpose.
- Answers one question per test. No decomposition across channels, creatives, and audiences the way MMM offers.
- Contamination risk. Media spillover across market boundaries, national campaigns that cannot be geo-split, and platform auto-optimization all bias results toward zero.
- Context-bound results. A test measures the effect in that season, at that spend level, against that competitive backdrop. It is a snapshot, not a constant. Retest periodically and at different budget levels.
- Requires genuine geographic targeting control, which some channels and some platforms do not give you.
- Easy to do badly. Two lines on a chart and an eyeballed gap are not a geo experiment. Without power analysis, pre-period validation, and placebo testing, the output is a number with unknown error bars.
Seven mistakes that ruin geo experiments
- No power analysis. Then reading a null result as evidence of no effect, when the design could never have detected one.
- Measuring platform-reported conversions instead of your own business KPI, which reintroduces the exact bias the experiment was meant to remove.
- Contaminated controls, where paused budget flows into control markets or a national campaign runs across both groups.
- Stopping early the moment the gap looks favourable.
- Ignoring carryover, by cutting measurement off the day spend stops.
- Testing during an atypical period, such as a peak-season week or a major competitive event.
- Treating one result as permanent truth. Incrementality changes with saturation, seasonality, and creative. Re-measure.
Tooling
Most credible tools implement variations on the same statistical ideas, so choose for fit with your data and team rather than for methodology novelty.
- GeoLift (Meta, open source, R): synthetic control based, with market selection and power calculators built in
- CausalImpact (Google, open source, R/Python): Bayesian structural time series for a single treated series
- Meridian (Google, open source): Bayesian MMM designed to accept experiment results as priors; Google has also announced Meridian GeoX for running the experiments themselves and feeding results back into the model
- Platform-native lift tests: convenient, but measured and reported by the media seller
- Custom implementations: worth it when your geo structure, KPI, or business model does not fit standard assumptions
The hard part was never the package. It is market selection, power, and defensible inference.
Frequently asked questions
How long should a geo experiment run? Typically four to eight weeks of treatment, plus a pre-period of at least a year for model fitting and a one to two week cool-down for carryover. Shorter tests are usually underpowered; much longer tests accumulate contamination risk from unrelated market changes.
How many geos do I need? As a rough floor, 20 to 30 assignable markets with meaningful volume. More is materially better. What matters more than raw count is whether any single market dominates your revenue, and how similar markets are to one another in the pre-period.
Are geo experiments better than MMM? They answer different questions. Geo experiments give a causal estimate for one intervention at a time. MMM covers every channel continuously but relies on modeling assumptions. The strongest measurement stacks run both and use experiments to calibrate the model.
Do I have to turn off spend to measure incrementality? No. Scale-up designs increase spend in treatment markets instead, which carries less business risk and measures marginal rather than average return. Holdouts remain the cleanest test of whether a channel contributes at all.
What if the result is not statistically significant? That is not the same as "no effect." Check the confidence interval against your power analysis. A null with an interval spanning -2% to +30% means the test was underpowered. A null with an interval of -3% to +4% is genuine evidence the effect is small.
Can geo experiments measure brand and upper-funnel media? Yes, and this is one of their strongest use cases, since no click-based method can. Expect to need longer windows, bigger spend deltas, and explicit carryover handling.
How much does a geo experiment cost? Two components: the analytics work, and the opportunity cost of the spend change. The second is usually the larger of the two, which is exactly why power analysis at the design stage matters so much.
How often should we re-run tests? Treat incrementality as a quantity that drifts. A program of two to four tests a year, rotating across channels and spend levels, keeps your MMM calibrated and your benchmarks honest.
Getting geo experiments right
Geo experiments are conceptually simple and easy to execute badly. The difference between a test that reshapes a budget and a test that produces a slide nobody trusts comes down to a handful of unglamorous decisions: whether the power analysis was run before the design was locked, whether the control group genuinely tracks the treatment group, whether the KPI came from your database or the platform's, and whether the decision rule was written down in advance.
Get those right and a single six-week test can change how millions of dollars are allocated, with a number you can actually defend.
Bell Statistics designs, runs and analyzes geo experiments for companies spending seriously across multiple markets, including market selection, power analysis, execution planning and MMM calibration. See how we work on geo testing →
Further reading from Bell:
Geo Testing at Bell Statistics
We plan, run and analyse geo experiments to measure what your marketing actually caused.
See how we work on geo testing →More on geo testing
All posts →
·3 min read
Geo Testing: Unlocking True Incrementality in Marketing and Product Experiments
Our “Geo Testing: Unlocking True Incrementality” webinar explored how teams can measure real-world impact when A/B testing isn’t possible - using geographic experiments and synthetic controls to reveal true incremental lift in marketing and product initiatives.

Allon Korem
Chief Executive Officer

·4 min read
Measuring the true effect of your ads with Geo Testing
Figuring out an ad's real effect is tricky. Clicks don't tell the whole story and attribution models fall short. The Solution: Geo Testing.

Amit Sasson
Causal Inference Expert

·5 min read
How to Tackle Your Marketing Challenges with Geo Tests
Traditional methods often fall short in measuring the true impact of marketing strategies. Here's how Geo Tests can resolve common marketing challenges.

Amit Sasson
Causal Inference Expert

