In plain English
A geo experiment randomises places rather than people, and the first design decision is how finely to cut the map. Postcodes, cities, designated market areas, regions and countries are all used. The choice looks administrative and it determines both how much statistical power the test can have and whether the estimate is biased, which makes it the most consequential decision in the design.
The tension is straightforward. Power comes from the number of units, because a geo test is cluster randomization with markets as clusters — twenty markets is twenty observations regardless of how many millions of people live in them. So finer units are better for precision. But advertising does not respect boundaries: television and radio buy at market level, outdoor reaches commuters from neighbouring areas, and people travel. Finer units mean more leakage between treated and untreated areas, which is interference and biases the effect towards zero.
The workable rule is to match the unit to how media is actually purchased and consumed. If your television is bought by DMA, randomising at postcode level guarantees contamination, because the buy itself crosses the boundaries you drew. If the campaign is digital and geo-targeted precisely, finer units become viable and the power gain is real. The unit should be the smallest area at which a treated and an untreated unit do not meaningfully share exposure.
The second constraint is data. A unit needs enough volume for its own outcome to be measurable — a market contributing three conversions a week adds noise rather than information, and a handful of large markets can end up dominating the estimate however many units you nominally have. Checking the distribution of volume across candidate units usually eliminates the finest option before any power calculation is done.
In practice most geo tests end up at DMA level in the US or at regional level elsewhere, with somewhere between twenty and a hundred usable units. That is a small sample by any other standard, which is why matching markets carefully on pre-period behaviour matters so much here and why the estimates carry wide intervals. Being honest about that width is part of reporting a geo test properly.
The formula
The two quantities that pull against each other, and the check that usually decides between candidate units.
- Power comes from unit count
effective n = number of geo unitsTwenty markets is twenty observations, whatever their population. This is the binding constraint in most geo designs.
- Leakage biases towards zero
observed ≈ true × ( 1 − spillover share )Finer units mean more cross-boundary exposure — see interference.
- The volume check
each unit needs enough conversions to be measurableCompute the distribution of weekly volume per candidate unit. Thin units add noise rather than information.
- The concentration check
share of total volume in the largest 3 unitsIf a few markets carry most of the outcome, the effective sample is smaller than the unit count — see the paired t-test calculator.
Worked example
A retailer plans a geo test for a television and digital campaign. Three granularities are compared on the same market: 210 DMAs, 51 states, and 1,400 usable postcode districts. Television is bought at DMA level.
- Postcode districts
- 1,400 units, median 18 conversions/week
- Postcode: TV spillover across boundaries
- substantial — TV is bought by DMA
- DMAs
- 210 units, median 340 conversions/week
- DMA: top 3 share of national volume
- 19%
- States
- 51 units, median 1,420 conversions/week
- Detectable effect: postcode / DMA / state
- biased / 6.1% / 12.4%
DMAs are the only viable choice: postcodes are contaminated by the television buy, and states give too few units to detect anything useful.
The postcode option has by far the best power on paper and is unusable, because the television buy crosses those boundaries by construction — treated and untreated postcodes inside one DMA see the same advertising, so the comparison is measuring a diluted version of the effect and there is no way to know by how much. States are clean and there are only 51 of them, giving a detectable effect of 12.4% that is larger than most campaigns produce. DMAs sit where the design has to sit: the unit matches how the media is bought, so contamination is limited, and 210 units support a 6.1% detectable effect. Note the concentration check passing at 19% — had three markets carried half the national volume, the effective sample would have been far below 210 and the DMA option would have been weaker than it looks.
Common misconceptions
- דFiner geo units are better because they give more statistical power.”
- They give more units and more leakage. If your media crosses the boundaries you drew — television bought by market, outdoor seen by commuters — treated and untreated units share exposure and the effect is attenuated by an unknown amount. Precision about a contaminated comparison is worth less than a wider interval around the right quantity.
- דA geo test with millions of people in each market has plenty of power.”
- Power comes from the number of markets, not the population inside them. Twenty markets behave like twenty observations however large they are, because everyone in a market receives the same treatment. Adding population to existing markets barely helps; adding markets does.
- דThe geo unit is an implementation detail the media team can decide.”
- It determines both the power ceiling and whether the estimate is biased, and those cannot be fixed in analysis. It does need the media team's input, because the right answer depends on how the buy is actually executed — which is exactly why it should be a joint decision rather than delegated.