
Measuring the true effect of your ads with Geo Testing
Figuring out an ad's real effect is tricky. Clicks don't tell the whole story and attribution models fall short. The Solution: Geo Testing.

A scale-up test adds spend to test markets; a scale-down test removes it. They are the same design run in opposite directions, and they answer different questions — one measures what more spend buys, the other what current spend is delivering.
A geo experiment can move spend in either direction. A scale-up test increases budget in the test markets and leaves control alone; a scale-down test cuts or removes it — a blackout or dark test when spend goes to zero. The mechanics are identical and the questions are not, so choosing the direction is choosing what you will learn.
Scale-down answers whether current spend is working. Switch a channel off in a set of markets and watch what happens to conversions: the gap is the incrementality of everything you were already doing. This is the right test for a channel you suspect is capturing demand rather than creating it — branded search being the standard case — and it is the only way to establish that a budget you are already committed to is earning its place.
Scale-up answers whether more spend would work, which is a different and usually harder question. Because of diminishing returns, an extra pound at current levels buys less than the average pound already spent, so the effect you are trying to detect is smaller than the average return. Scale-up tests are correspondingly less sensitive and need either larger budget increases or more markets to resolve anything.
The asymmetry in sensitivity is worth being explicit about. Going from full spend to zero produces the largest possible signal, so a blackout test detects an effect with fewer markets or a shorter run than any scale-up variant. The cost is commercial rather than statistical: you are deliberately forgoing revenue in the dark markets for the duration, and that has to be authorised by someone who understands what is being spent to learn.
In practice the sequence that works is scale-down first, then scale-up. Establish that current spend is incremental at all, since a channel with near-zero incrementality should be cut rather than optimised, and only then test whether increasing it pays. Running scale-up on a channel whose baseline contribution has never been established risks measuring the marginal return on something that was never returning anything.
Both directions estimate the same kind of quantity; the difference is which part of the response curve they sit on.
incremental value of current spendCompares full spend against zero (or reduced). The largest available signal.
marginal return on additional spendCompares current against increased. Smaller by construction because of diminishing returns.
Δspend is larger, and marginal return < average returnBoth terms favour the downward direction — see diminishing returns.
forgone revenue = incremental value × dark markets × durationReal money, and the point of the test is that you do not yet know the first term — see the sample size calculator.
An advertiser wants to know whether its branded search spend is worth keeping and whether its social budget should grow. Both questions are put to geo tests, one in each direction, across 60 matched markets.
Switching branded search off cost 1.8% of conversions and saved £182,000. Increasing social by 60% produced an effect that cannot be distinguished from zero.
The two results illustrate the asymmetry directly. The scale-down gives a clean answer: branded search is doing something, but only 1.8% of conversions for the money, which makes the arithmetic of keeping it at that level questionable and worth a proper iROAS calculation. The scale-up is inconclusive, and note that this is the expected outcome rather than a failure — a 60% budget increase sits on the flat part of the response curve, so the marginal effect is small and 30 markets over six weeks cannot resolve it. The honest conclusion is not "social does not work" but "this test could not detect the marginal return", and the options are a larger budget swing, more markets, or accepting that the question needs a scale-down instead. Running the branded search test first was the right order: had it shown near-zero incrementality, the channel would have been a cut rather than an optimisation target.

Figuring out an ad's real effect is tricky. Clicks don't tell the whole story and attribution models fall short. The Solution: Geo Testing.


Traditional methods often fall short in measuring the true impact of marketing strategies. Here's how Geo Tests can resolve common marketing challenges.

Applying it to a live measurement problem is the part that goes wrong. If you are designing an experiment, reading a result you do not trust, or trying to work out what your marketing actually caused, that is the work we do.