Bell Statistics

Scale-up or scale-down geo test?

A scale-up test adds spend to test markets; a scale-down test removes it. They are the same design run in opposite directions, and they answer different questions — one measures what more spend buys, the other what current spend is delivering.

Also called
scale-up test, scale-down test, blackout test, dark test, spend-down test, holdout market test
Allon Korem

Written by Allon Korem

Chief Executive Officer

Last updated

In plain English

A geo experiment can move spend in either direction. A scale-up test increases budget in the test markets and leaves control alone; a scale-down test cuts or removes it — a blackout or dark test when spend goes to zero. The mechanics are identical and the questions are not, so choosing the direction is choosing what you will learn.

Scale-down answers whether current spend is working. Switch a channel off in a set of markets and watch what happens to conversions: the gap is the incrementality of everything you were already doing. This is the right test for a channel you suspect is capturing demand rather than creating it — branded search being the standard case — and it is the only way to establish that a budget you are already committed to is earning its place.

Scale-up answers whether more spend would work, which is a different and usually harder question. Because of diminishing returns, an extra pound at current levels buys less than the average pound already spent, so the effect you are trying to detect is smaller than the average return. Scale-up tests are correspondingly less sensitive and need either larger budget increases or more markets to resolve anything.

The asymmetry in sensitivity is worth being explicit about. Going from full spend to zero produces the largest possible signal, so a blackout test detects an effect with fewer markets or a shorter run than any scale-up variant. The cost is commercial rather than statistical: you are deliberately forgoing revenue in the dark markets for the duration, and that has to be authorised by someone who understands what is being spent to learn.

In practice the sequence that works is scale-down first, then scale-up. Establish that current spend is incremental at all, since a channel with near-zero incrementality should be cut rather than optimised, and only then test whether increasing it pays. Running scale-up on a channel whose baseline contribution has never been established risks measuring the marginal return on something that was never returning anything.

The formula

Both directions estimate the same kind of quantity; the difference is which part of the response curve they sit on.

Scale-down estimand
incremental value of current spend

Compares full spend against zero (or reduced). The largest available signal.

Scale-up estimand
marginal return on additional spend

Compares current against increased. Smaller by construction because of diminishing returns.

Why scale-down is more sensitive
Δspend is larger, and marginal return < average return

Both terms favour the downward direction — see diminishing returns.

The cost of a blackout
forgone revenue = incremental value × dark markets × duration

Real money, and the point of the test is that you do not yet know the first term — see the sample size calculator.

Worked example

An advertiser wants to know whether its branded search spend is worth keeping and whether its social budget should grow. Both questions are put to geo tests, one in each direction, across 60 matched markets.

Branded search: scale-down (to zero) in 30 markets
4 weeks
Branded search: conversions in dark markets
−1.8% vs control
Branded search: spend saved
£182,000 over the period
Social: scale-up (+60% budget) in 30 markets
6 weeks
Social: conversions in boosted markets
+2.1%, 95% CI −0.4% to +4.6%
Social: additional spend
£240,000

Switching branded search off cost 1.8% of conversions and saved £182,000. Increasing social by 60% produced an effect that cannot be distinguished from zero.

The two results illustrate the asymmetry directly. The scale-down gives a clean answer: branded search is doing something, but only 1.8% of conversions for the money, which makes the arithmetic of keeping it at that level questionable and worth a proper iROAS calculation. The scale-up is inconclusive, and note that this is the expected outcome rather than a failure — a 60% budget increase sits on the flat part of the response curve, so the marginal effect is small and 30 markets over six weeks cannot resolve it. The honest conclusion is not "social does not work" but "this test could not detect the marginal return", and the options are a larger budget swing, more markets, or accepting that the question needs a scale-down instead. Running the branded search test first was the right order: had it shown near-zero incrementality, the channel would have been a cut rather than an optimisation target.

Common misconceptions

Scale-up and scale-down tests measure the same thing in opposite directions.
They measure different points on the response curve. Scale-down gives the incremental value of spend you are already making; scale-up gives the marginal return on additional spend, which is smaller because of diminishing returns. A channel can be clearly incremental at current levels and produce nothing extra when increased.
A blackout test is too risky to run.
It is the most sensitive design available and its cost is bounded and calculable — spend forgone in the dark markets for the duration, which is a fraction of national revenue. The genuinely expensive option is continuing to fund a channel whose incremental contribution has never been established, which is often a much larger number.
An inconclusive scale-up test shows the channel does not scale.
More often it shows the test lacked the sensitivity to detect a marginal effect. The signal from a budget increase is small by construction, so an inconclusive result is the default outcome unless the swing is large and the market count is high. Distinguish 'no effect detected' from 'no effect' by looking at the interval.

Frequently asked questions

Which direction should I run first?
Scale-down, in almost every case. It establishes whether current spend is incremental at all, which is the prior question — a channel contributing near zero should be cut rather than optimised, and testing whether to increase it would be measuring the marginal return on something that was not returning anything. Once the baseline contribution is established, scale-up answers whether growth pays.
How long should a blackout test run?
Long enough to cover the conversion lag plus a stable measurement window, which for most consumer purchases means three to six weeks. Too short and delayed conversions from before the blackout mask the effect; too long and the forgone revenue mounts while brand effects start to accumulate. Add a cooldown period after the blackout ends to observe recovery, which also tests whether the effect was demand shifted rather than lost.
How big does a scale-up budget increase need to be?
Large enough that the marginal effect is detectable, which usually means 50% or more rather than a 10% increment. Because the marginal return is smaller than the average return, a modest increase produces a signal well below what a few dozen markets can resolve. If the budget for a large swing is not available, a scale-down test on the existing spend will answer a related question far more cheaply.

Related terms

  • Geo test periods

    Match, measure, then wait — and the cooldown is the phase teams skip and then misread the result.

  • Geo unit

    How finely you cut the map: more units means more power, and more spillover between them.

  • Incrementality

    The conversions that would not have happened anyway — and the gap between that and what platforms report.

  • Test and control markets

    Too few markets for randomisation to balance them, so they are matched — and the matching is the whole design.

  • Budget scaling

    How much to move, and what the next pound buys — which is never what the last pound averaged.

Calculate it

  • A/B test sample size

    Size a two-proportion experiment before you launch, then read the lift and its interval once it lands.

  • Two-sample t-test

    Compare the average of two independent groups — plan the sample size, then test the result.

Knowing the term is the easy part

Applying it to a live measurement problem is the part that goes wrong. If you are designing an experiment, reading a result you do not trust, or trying to work out what your marketing actually caused, that is the work we do.