Bell Statistics

What is a geo unit?

A geo unit is the geographic area a geo experiment randomises — a postcode, a city, a media market or a whole country. The choice sets how many units you have and how much advertising leaks between them, and those two pull in opposite directions.

Also called
geographic unit, geo granularity, market unit, geo cell
Allon Korem

Written by Allon Korem

Chief Executive Officer

Last updated

In plain English

A geo experiment randomises places rather than people, and the first design decision is how finely to cut the map. Postcodes, cities, designated market areas, regions and countries are all used. The choice looks administrative and it determines both how much statistical power the test can have and whether the estimate is biased, which makes it the most consequential decision in the design.

The tension is straightforward. Power comes from the number of units, because a geo test is cluster randomization with markets as clusters — twenty markets is twenty observations regardless of how many millions of people live in them. So finer units are better for precision. But advertising does not respect boundaries: television and radio buy at market level, outdoor reaches commuters from neighbouring areas, and people travel. Finer units mean more leakage between treated and untreated areas, which is interference and biases the effect towards zero.

The workable rule is to match the unit to how media is actually purchased and consumed. If your television is bought by DMA, randomising at postcode level guarantees contamination, because the buy itself crosses the boundaries you drew. If the campaign is digital and geo-targeted precisely, finer units become viable and the power gain is real. The unit should be the smallest area at which a treated and an untreated unit do not meaningfully share exposure.

The second constraint is data. A unit needs enough volume for its own outcome to be measurable — a market contributing three conversions a week adds noise rather than information, and a handful of large markets can end up dominating the estimate however many units you nominally have. Checking the distribution of volume across candidate units usually eliminates the finest option before any power calculation is done.

In practice most geo tests end up at DMA level in the US or at regional level elsewhere, with somewhere between twenty and a hundred usable units. That is a small sample by any other standard, which is why matching markets carefully on pre-period behaviour matters so much here and why the estimates carry wide intervals. Being honest about that width is part of reporting a geo test properly.

The formula

The two quantities that pull against each other, and the check that usually decides between candidate units.

Power comes from unit count
effective n = number of geo units

Twenty markets is twenty observations, whatever their population. This is the binding constraint in most geo designs.

Leakage biases towards zero
observed ≈ true × ( 1 − spillover share )

Finer units mean more cross-boundary exposure — see interference.

The volume check
each unit needs enough conversions to be measurable

Compute the distribution of weekly volume per candidate unit. Thin units add noise rather than information.

The concentration check
share of total volume in the largest 3 units

If a few markets carry most of the outcome, the effective sample is smaller than the unit count — see the paired t-test calculator.

Worked example

A retailer plans a geo test for a television and digital campaign. Three granularities are compared on the same market: 210 DMAs, 51 states, and 1,400 usable postcode districts. Television is bought at DMA level.

Postcode districts
1,400 units, median 18 conversions/week
Postcode: TV spillover across boundaries
substantial — TV is bought by DMA
DMAs
210 units, median 340 conversions/week
DMA: top 3 share of national volume
19%
States
51 units, median 1,420 conversions/week
Detectable effect: postcode / DMA / state
biased / 6.1% / 12.4%

DMAs are the only viable choice: postcodes are contaminated by the television buy, and states give too few units to detect anything useful.

The postcode option has by far the best power on paper and is unusable, because the television buy crosses those boundaries by construction — treated and untreated postcodes inside one DMA see the same advertising, so the comparison is measuring a diluted version of the effect and there is no way to know by how much. States are clean and there are only 51 of them, giving a detectable effect of 12.4% that is larger than most campaigns produce. DMAs sit where the design has to sit: the unit matches how the media is bought, so contamination is limited, and 210 units support a 6.1% detectable effect. Note the concentration check passing at 19% — had three markets carried half the national volume, the effective sample would have been far below 210 and the DMA option would have been weaker than it looks.

Common misconceptions

Finer geo units are better because they give more statistical power.
They give more units and more leakage. If your media crosses the boundaries you drew — television bought by market, outdoor seen by commuters — treated and untreated units share exposure and the effect is attenuated by an unknown amount. Precision about a contaminated comparison is worth less than a wider interval around the right quantity.
A geo test with millions of people in each market has plenty of power.
Power comes from the number of markets, not the population inside them. Twenty markets behave like twenty observations however large they are, because everyone in a market receives the same treatment. Adding population to existing markets barely helps; adding markets does.
The geo unit is an implementation detail the media team can decide.
It determines both the power ceiling and whether the estimate is biased, and those cannot be fixed in analysis. It does need the media team's input, because the right answer depends on how the buy is actually executed — which is exactly why it should be a joint decision rather than delegated.

Frequently asked questions

How do I choose the geo unit?
Start from how the media is bought and consumed: the unit should be the smallest area where a treated and an untreated unit do not meaningfully share exposure. Then check volume — each unit needs enough conversions to be measurable — and check concentration, because a few dominant markets shrink the effective sample below the unit count. Those three checks usually leave one viable option.
How many geo units do I need?
Enough that the number of units, not the population, supports the effect you want to detect. Twenty is very few and will only resolve large effects; fifty to a couple of hundred is the usual workable range. Run the power calculation on the unit count with the observed market-to-market variability, and expect the answer to be less encouraging than a user-level calculation would suggest.
Can a geo test use whole countries as units?
Only with many countries, and it brings problems beyond the small sample. Countries differ in seasonality, competitive landscape, currency and consumer behaviour, so matching them on pre-period trends is much harder than matching cities within one market. Where a campaign genuinely runs at country level, synthetic control with a carefully constructed donor pool is usually a better fit than a straightforward test-and-control split.

Related terms

  • Cluster randomization

    Assign the group, not the person — the remedy for interference, paid for in statistical power.

  • Designated Market Area

    The 210 US television markets — the standard geo unit because the media buy already respects those boundaries.

  • Scale-up vs scale-down test

    Add budget or switch it off — the direction decides which question you get an answer to.

  • Test and control markets

    Too few markets for randomisation to balance them, so they are matched — and the matching is the whole design.

Calculate it

  • A/B test sample size

    Size a two-proportion experiment before you launch, then read the lift and its interval once it lands.

  • Paired t-test

    Before-and-after or matched pairs — size the study on the difference SD, then test it.

Knowing the term is the easy part

Applying it to a live measurement problem is the part that goes wrong. If you are designing an experiment, reading a result you do not trust, or trying to work out what your marketing actually caused, that is the work we do.

References

  • Vaver, J., & Koehler, J. (2011). Measuring Ad Effectiveness Using Geo Experiments. Google Inc.
  • Kerman, J., Wang, P., & Vaver, J. (2017). Estimating Ad Effectiveness using Geo Experiments in a Time-Based Regression Framework. Google Inc.