Bell Statistics

What is a geo experiment?

A geo experiment randomises whole geographic regions rather than individual users, running marketing in some and withholding it in others. It recovers a genuine randomised comparison for anything that cannot be hidden from an individual person.

Also called
geo test, geo lift test, matched market test
Allon Korem

Written by Allon Korem

Chief Executive Officer

Last updated

In plain English

A/B testing needs to show different things to different people, which rules it out for most marketing. You cannot show a television advertisement to half a household. You cannot hide a billboard, a radio spot, or a price change from a randomly chosen cookie. And even where you technically can — social and display — a user seen on three devices and two browsers is not one unit. A geo experiment sidesteps all of this by moving the randomisation up a level: regions are assigned to treatment or control, spend is set accordingly, and outcomes are compared per head of population.

What you buy is a real causal estimate for channels that otherwise have none. Because assignment is random, the treatment and control regions are comparable in expectation on everything — including local demand, competitive intensity and the seasonal patterns that make observational marketing analysis so treacherous. That makes a geo test the most direct available measurement of incrementality: the difference between regions is what the advertising caused, not what it was present for.

What you pay is statistical power. A country has perhaps 50 to 210 usable regions, so the effective sample size is the number of regions rather than the number of people, and regions vary enormously in size and character. This is why geo tests detect double-digit lifts comfortably and struggle below about 5%, and why the design work matters more than in a user-level test: stratify regions into blocks of similar size and baseline before randomising, or match pairs and randomise within them. A test that assigns the three largest metros to treatment by chance has burnt most of its precision before collecting a single data point.

The analysis is rarely a simple difference in means. Regions differ in level, so the standard approaches are difference-in-differences against a pre-period, or a synthetic control per treated region — and the second is preferred when regions were not randomised, which happens more often than anyone admits. Standard errors must be clustered at the region, since outcomes within a region across weeks are highly correlated, and ignoring that produces intervals several times too narrow.

Three practical constraints decide whether a geo test is feasible. Spillover: regions must be far enough apart that treated advertising does not reach control areas, which rules out adjacent metros for broadcast media and makes online targeting accuracy a real concern. Duration: four to eight weeks is typical, long enough for effects to accumulate and short enough that the regions do not drift apart for unrelated reasons. And a clean pre-period of several months, because almost every analysis method leans on it. We cover the design decisions in how to tackle your marketing challenges with geo tests.

The formula

The estimator is ordinary; the sample size arithmetic is where geo tests differ from user-level ones, because n is the number of regions.

Per-capita outcome
Yᵢ = revenueᵢ / populationᵢ

Normalising by population is what makes regions of wildly different size comparable. Weighting by population afterwards recovers the aggregate effect.

Number of regions needed
k = 2·( z₁₋α/₂ + z₁₋β )² · σ²_region / δ²

σ² is the variance ACROSS regions, not across people. This is why a country with 60 usable regions supports a 10% test and not a 3% one — see the two-sample t-test calculator.

Difference-in-differences estimate
τ̂ = ( Ȳ_t,post − Ȳ_t,pre ) − ( Ȳ_c,post − Ȳ_c,pre )

Removes fixed differences between regions and anything that affected all of them over the window.

Incremental return
iROAS = ( τ̂ × population_treated ) / incremental spend

The number the budget decision needs. Use the interval as well as the point estimate — geo tests are noisy enough that the two often imply different decisions.

Worked example

An advertiser wants to know whether a connected-TV campaign is worth £1.2m a quarter. There are 68 usable metro areas. They stratify into 34 pairs matched on population and baseline revenue per head, randomise one of each pair into treatment, and run for six weeks with twelve weeks of pre-period.

Regions
34 treated, 34 control
Design
Matched pairs, randomised within pair
Incremental spend
£600,000 over six weeks
DiD estimate
+£1.94 revenue per head
Clustered 95% CI
+£0.61 to +£3.27
Treated population
1.42 million

Incremental revenue of about £2.76m against £600,000 of spend — an iROAS of 4.6, with the interval running from 1.4 to 7.7.

The point estimate says the campaign returns £4.60 per pound, and the interval says it returns somewhere between £1.40 and £7.70. If the business needs 2.0 to justify the channel, the honest summary is that it probably clears the bar but the test could not confirm it — and that width is the characteristic cost of geo testing, where 68 regions is the sample size no matter how many millions of people they contain. Two design choices did the heavy lifting: matched pairs prevented a randomisation that put the big metros on one side, and clustering at the region kept the interval honest. Analysing the same data at week level without clustering would have produced an interval perhaps a third as wide and a false sense of precision.

Common misconceptions

We have millions of users in the test, so it is well powered.
The unit of randomisation is the region, so the effective sample size is the number of regions — often fewer than a hundred. Millions of users spread across sixty regions give you sixty independent observations, not millions, and analysing at user level without clustering produces intervals that are badly too narrow.
We picked comparable markets rather than randomising, which is cleaner.
Choosing markets reintroduces exactly the selection the design exists to remove, because the reasons a market looks suitable are often related to its trajectory. Randomise, ideally within matched strata so you get balance and randomness together. If markets were already chosen, analyse with synthetic control and state the assumption.
A geo test can measure any channel.
It needs the treatment to stay inside its region. Broadcast media bleeds across adjacent metros, online targeting is imperfect near boundaries, and people travel — all of which contaminate control regions and bias the effect towards zero. Choose geographically separated regions and expect the design to be unusable for some channels.

Frequently asked questions

How many regions do I need for a geo experiment?
Enough that the variance across regions supports the effect you want to detect, which usually means at least 20 per arm and preferably more. The binding constraint is that regions are the sample, so a country with 60 usable ones supports detecting a 10% lift comfortably and a 3% lift not at all. Matched-pair randomisation and per-capita normalisation both help materially, and are worth more than adding a handful of marginal regions.
How long should a geo test run?
Four to eight weeks in most cases, plus several months of clean pre-period for the analysis to lean on. Shorter than four weeks rarely lets brand effects accumulate or lets you distinguish signal from weekly noise across regions. Much longer and the regions drift apart for reasons unrelated to the test — a store opening, a local event, a competitor's push — which the design cannot separate from the treatment.
Should I run a geo test or build a marketing mix model?
They answer different questions and work well together. A geo test measures one channel or one campaign with a design-based causal claim, precisely and expensively. A mix model estimates all channels at once from historical variation, cheaply and with weaker identification. The strongest arrangement is to run geo tests periodically and use their results to calibrate the model's priors, which anchors the model on something that was actually randomised.
How do I handle spillover between regions?
Design around it first: choose regions that are geographically separated, avoid assigning adjacent metros to different arms, and exclude buffer zones near boundaries from the analysis. Where spillover remains, it biases the measured effect towards zero, since control regions receive some treatment — so a positive result is conservative, and a null one is genuinely ambiguous between no effect and heavy contamination.

Related terms

  • Difference-in-differences

    Subtract the untreated group's change from the treated group's — and everything rests on parallel trends.

  • Incrementality

    The conversions that would not have happened anyway — and the gap between that and what platforms report.

  • Marketing mix modelling

    One regression across every channel, built on aggregate data — no tracking, and strong assumptions.

  • Randomization

    The one mechanism that makes a comparison causal — and the four ways it silently fails.

  • Synthetic control

    Build the comparison group instead of finding one — the method for when you have one treated unit.

Calculate it

  • Two-sample t-test

    Compare the average of two independent groups — plan the sample size, then test the result.

  • A/B test sample size

    Size a two-proportion experiment before you launch, then read the lift and its interval once it lands.

  • Paired t-test

    Before-and-after or matched pairs — size the study on the difference SD, then test it.

Knowing the term is the easy part

Region selection, matching and power are where geo tests are won or lost, long before the analysis — see Geo Testing

References

  • Vaver, J., & Koehler, J. (2011). Measuring Ad Effectiveness Using Geo Experiments. Google Inc. Technical Report.
  • Kerman, J., Wang, P., & Vaver, J. (2017). Estimating Ad Effectiveness using Geo Experiments in a Time-Based Regression Framework. Google Inc. Technical Report.