Bell Statistics

How are test markets chosen?

Test and control markets are the two sets of geographies a geo experiment compares. Because there are usually only a few dozen units, they are matched on pre-period behaviour rather than split at random — the matching does the work randomisation cannot at that sample size.

Also called
matched markets, matched market test, test markets, control markets, market matching
Allon Korem

Written by Allon Korem

Chief Executive Officer

Last updated

In plain English

A geo experiment divides markets into a treated set that receives the campaign and a control set that does not. With user-level experiments, randomisation handles balance automatically because the sample is enormous. With twenty or eighty markets it does not — a random split can easily put most of the large or fast-growing markets on one side, and at that sample size such an imbalance is likely rather than unlucky.

So geo tests match. The standard approach pairs or groups markets on their pre-period behaviour: markets whose weekly sales tracked each other closely before the test are assumed to continue tracking each other during it, so the control set forms a credible counterfactual for the treated one. The correlation between the two sets over the pre-period is the single number that tells you whether the design is sound, and it should be checked before any spend is committed.

Matching is not a substitute for randomisation and it makes a different assumption. Randomisation balances everything, including what you did not measure. Matching balances only what you matched on, so a market with an unobserved characteristic — a competitor about to launch there, a store closing — remains a threat. The usual compromise is to match into pairs or strata and then randomise within them, which keeps the balance and restores some of randomisation's protection against the unobserved.

The practical mechanics matter more here than in a user-level test because there is so little data to recover from. Exclude markets that are structurally unusual before matching rather than after seeing results. Match on the outcome metric itself over a period long enough to cover seasonality, not on population. And check the matched pairs visually — a pair with high correlation driven by one shared spike is not really matched, and the number alone will not show that.

The threat that ends geo tests is a matched pair diverging for a reason unrelated to the campaign. A store closure, a local competitor promotion, a weather event: any of these breaks the assumption that the control market predicts the treated one, and with only a few dozen units a single broken pair can move the estimate materially. Monitoring the pairs during the test, rather than only at the end, is what allows a broken pair to be identified and handled honestly.

The formula

The matching criterion, the check that matters most, and the estimator the design supports.

The matching criterion
maximise correlation of the outcome across the pre-period

Match on the metric you will measure, over a window long enough to include seasonality.

The design check
ρ( test aggregate, control aggregate ) over the pre-period

Above about 0.9 is workable. Below that, the control set is a weak counterfactual — see the correlation calculator.

Match then randomise
form pairs on pre-period behaviour, randomise within each pair

Keeps the balance and restores some protection against unobserved differences.

The estimator
paired comparison of within-pair differences

Removes between-market variation entirely — see the paired t-test calculator.

Worked example

A retailer selects 40 markets for a geo test, forming 20 matched pairs on 26 weeks of pre-period revenue and randomising within each pair. During the test, one pair diverges sharply for a reason unrelated to the campaign.

Markets
40, formed into 20 matched pairs
Pre-period correlation, aggregate
0.96
Weakest pair correlation
0.71
Estimated lift, all 20 pairs
+3.9%, 95% CI +0.2% to +7.6%
Pair 14: control market
competitor opened two stores mid-test
Estimated lift, excluding pair 14
+2.6%, 95% CI +0.6% to +4.6%

One broken pair out of twenty moved the estimate by 1.3 percentage points and widened the interval substantially.

The influence of a single pair is the thing to internalise about geo tests. With twenty pairs each carries 5% of the weight, and a pair where the control market was disrupted contributes a spurious difference that is indistinguishable from campaign effect in the aggregate. Excluding it is defensible here — the competitor opening is documented, unrelated to the campaign, and was identified from the pair-level monitoring rather than by looking for the exclusion that improved the result. That distinction is everything: the same exclusion decided after seeing which pair was inconvenient would be indefensible. Note the weakest pair correlation of 0.71 was visible before the test started and should have prompted either a better match or dropping that pair at design time. Checking the pairs individually rather than only the aggregate 0.96 is what would have caught it.

Common misconceptions

Matching markets is just a version of randomisation.
Randomisation balances everything including unobserved characteristics; matching balances only what you matched on. A market with an unmeasured difference — an imminent competitor launch, a store closure — stays unbalanced. Matching into pairs and then randomising within them is the design that gets some of both.
A high aggregate pre-period correlation means the design is sound.
The aggregate can look excellent while individual pairs are poorly matched, and with twenty pairs a single bad one carries 5% of the weight. Check the pairs individually and look at them, not just their correlation coefficients — a pair whose correlation comes from one shared seasonal spike is not matched in any useful sense.
Markets should be matched on population or size.
Match on the outcome metric over the pre-period, which is what has to keep tracking during the test. Two markets of identical population can have completely different sales trajectories, and two of very different sizes can move in near-perfect proportion. Size matters for weighting the analysis, not for choosing the pairs.

Frequently asked questions

What should markets be matched on?
The outcome metric itself, over a pre-period long enough to cover a full seasonal cycle — typically six months to a year. Population and demographics are poor proxies, because two similarly sized markets can behave completely differently. After matching, inspect the pairs visually rather than trusting the correlation coefficient alone; a pair whose agreement comes from one shared spike is not usefully matched.
What do I do if a matched pair breaks during the test?
Identify it from pair-level monitoring during the test rather than at the end, document the external cause, and decide before looking at the effect whether it warrants exclusion. That ordering is what separates a defensible exclusion from choosing the answer. Report the estimate both with and without the pair, so a reader can see how much rests on that decision.
How many market pairs do I need?
Twenty pairs is a common minimum and is genuinely few — each carries 5% of the weight, so one anomaly moves the result materially. Forty or more is considerably safer where the market count allows. The relevant power calculation runs on the number of pairs and the variability of within-pair differences, not on the population inside the markets.

Related terms

  • Designated Market Area

    The 210 US television markets — the standard geo unit because the media buy already respects those boundaries.

  • Geo unit

    How finely you cut the map: more units means more power, and more spillover between them.

  • Scale-up vs scale-down test

    Add budget or switch it off — the direction decides which question you get an answer to.

  • Synthetic control

    Build the comparison group instead of finding one — the method for when you have one treated unit.

  • GeoLift

    Open-source geo testing with the power simulation built in — it tells you whether the test can work before you run it.

Calculate it

  • Paired t-test

    Before-and-after or matched pairs — size the study on the difference SD, then test it.

  • Correlation test

    Pearson r or Spearman rho, with the Fisher-z interval that says how little a small sample knows.

Knowing the term is the easy part

Applying it to a live measurement problem is the part that goes wrong. If you are designing an experiment, reading a result you do not trust, or trying to work out what your marketing actually caused, that is the work we do.