
How to Tackle Your Marketing Challenges with Geo Tests
Traditional methods often fall short in measuring the true impact of marketing strategies. Here's how Geo Tests can resolve common marketing challenges.

Test and control markets are the two sets of geographies a geo experiment compares. Because there are usually only a few dozen units, they are matched on pre-period behaviour rather than split at random — the matching does the work randomisation cannot at that sample size.
A geo experiment divides markets into a treated set that receives the campaign and a control set that does not. With user-level experiments, randomisation handles balance automatically because the sample is enormous. With twenty or eighty markets it does not — a random split can easily put most of the large or fast-growing markets on one side, and at that sample size such an imbalance is likely rather than unlucky.
So geo tests match. The standard approach pairs or groups markets on their pre-period behaviour: markets whose weekly sales tracked each other closely before the test are assumed to continue tracking each other during it, so the control set forms a credible counterfactual for the treated one. The correlation between the two sets over the pre-period is the single number that tells you whether the design is sound, and it should be checked before any spend is committed.
Matching is not a substitute for randomisation and it makes a different assumption. Randomisation balances everything, including what you did not measure. Matching balances only what you matched on, so a market with an unobserved characteristic — a competitor about to launch there, a store closing — remains a threat. The usual compromise is to match into pairs or strata and then randomise within them, which keeps the balance and restores some of randomisation's protection against the unobserved.
The practical mechanics matter more here than in a user-level test because there is so little data to recover from. Exclude markets that are structurally unusual before matching rather than after seeing results. Match on the outcome metric itself over a period long enough to cover seasonality, not on population. And check the matched pairs visually — a pair with high correlation driven by one shared spike is not really matched, and the number alone will not show that.
The threat that ends geo tests is a matched pair diverging for a reason unrelated to the campaign. A store closure, a local competitor promotion, a weather event: any of these breaks the assumption that the control market predicts the treated one, and with only a few dozen units a single broken pair can move the estimate materially. Monitoring the pairs during the test, rather than only at the end, is what allows a broken pair to be identified and handled honestly.
The matching criterion, the check that matters most, and the estimator the design supports.
maximise correlation of the outcome across the pre-periodMatch on the metric you will measure, over a window long enough to include seasonality.
ρ( test aggregate, control aggregate ) over the pre-periodAbove about 0.9 is workable. Below that, the control set is a weak counterfactual — see the correlation calculator.
form pairs on pre-period behaviour, randomise within each pairKeeps the balance and restores some protection against unobserved differences.
paired comparison of within-pair differencesRemoves between-market variation entirely — see the paired t-test calculator.
A retailer selects 40 markets for a geo test, forming 20 matched pairs on 26 weeks of pre-period revenue and randomising within each pair. During the test, one pair diverges sharply for a reason unrelated to the campaign.
One broken pair out of twenty moved the estimate by 1.3 percentage points and widened the interval substantially.
The influence of a single pair is the thing to internalise about geo tests. With twenty pairs each carries 5% of the weight, and a pair where the control market was disrupted contributes a spurious difference that is indistinguishable from campaign effect in the aggregate. Excluding it is defensible here — the competitor opening is documented, unrelated to the campaign, and was identified from the pair-level monitoring rather than by looking for the exclusion that improved the result. That distinction is everything: the same exclusion decided after seeing which pair was inconvenient would be indefensible. Note the weakest pair correlation of 0.71 was visible before the test started and should have prompted either a better match or dropping that pair at design time. Checking the pairs individually rather than only the aggregate 0.96 is what would have caught it.

Traditional methods often fall short in measuring the true impact of marketing strategies. Here's how Geo Tests can resolve common marketing challenges.


Figuring out an ad's real effect is tricky. Clicks don't tell the whole story and attribution models fall short. The Solution: Geo Testing.

Applying it to a live measurement problem is the part that goes wrong. If you are designing an experiment, reading a result you do not trust, or trying to work out what your marketing actually caused, that is the work we do.