Bell Statistics

What are the geo test periods?

A geo experiment runs in three phases: a pre-period used to match the markets and establish the baseline relationship between them, a test period during which spend differs between the arms, and a cooldown afterwards that captures delayed and borrowed-forward effects.

Also called
pre-period, test period, cooldown period, washout period, post-period
Allon Korem

Written by Allon Korem

Chief Executive Officer

Last updated

In plain English

A geo experiment has a timeline rather than just a duration, and each phase does a distinct job. The pre-period establishes how the markets relate to each other before anything changes. The test period is when spend differs between arms. The cooldown follows the test and captures effects that arrive late. Getting any of the three wrong produces a number that looks fine and answers the wrong question.

The pre-period does double duty and is the phase that decides whether the design is sound at all. It supplies the data for matching test and control markets, and it establishes the baseline relationship the estimate depends on — the correlation between the two sets over that window is the single best indicator of whether the control set is a credible counterfactual. It needs to be long enough to cover a full seasonal cycle, which for most businesses means six months to a year rather than the few weeks teams often allocate.

The test period has to be long enough for the media to work and short enough to afford. The constraint people underestimate is the conversion lag: if purchases typically happen two weeks after exposure, a four-week test spends its first fortnight accumulating conversions from pre-test exposure. That dilutes the measured effect, and the usual remedy is to discard an initial window from the analysis rather than to extend the test indefinitely.

The cooldown is the phase most often skipped and it is what separates a real effect from a timing artefact. A campaign that pulls purchases forward looks identical to one that creates them during the test period; only the weeks afterwards distinguish them. If the treated markets dip below control after the campaign ends, demand was borrowed rather than made, and the test-period lift overstates the true contribution — sometimes to the point of reversing the conclusion.

The cooldown also matters for the next test on the same markets. Adstock means advertising keeps working after it stops, so running a new experiment immediately means the previous campaign's residual effect is still present in the markets that received it. A washout of several weeks between tests is what stops one experiment contaminating the next, and it is a real constraint on how many geo tests a year a business can actually run.

The formula

Each phase has a length driven by something specific, and getting the driver right matters more than any rule of thumb.

Pre-period
at least one full seasonal cycle

Six to twelve months. Used for matching and for establishing the baseline correlation.

Test period
conversion lag + a stable measurement window

Discard the initial lag window from the analysis rather than extending the whole test.

Cooldown
2 to 3 × the adstock half-life

Long enough to see whether treated markets dip below control — see adstock.

The design check
ρ( test, control ) across the pre-period

Above about 0.9 for a workable design — see the correlation calculator.

Worked example

A four-week geo campaign is measured with a two-week cooldown that the team almost skipped. Weekly conversions in treated markets are compared against control across all three phases.

Pre-period
26 weeks, test/control correlation 0.95
Test weeks 1-2
+1.1% — diluted by pre-test conversion lag
Test weeks 3-4
+7.4%
Test period overall
+4.3%
Cooldown weeks 1-2
−3.1% below control
Net effect across test + cooldown
+1.9%

The test period shows +4.3%. Including the cooldown, the net effect is +1.9% — well under half, because much of the lift was pulled forward.

Two phases are doing corrective work here and both would have been missed by a bare test-period readout. The first two weeks were diluted by conversions arriving from pre-test exposure, which is why weeks 3-4 show a much larger effect — the honest test-period estimate discards that initial window rather than averaging it in. More importantly, the cooldown shows treated markets running 3.1% below control after the campaign ended, which is the signature of borrowed demand: customers who would have bought in weeks 5-6 bought in weeks 3-4 instead. The campaign created some genuine incremental demand, and less than half what the test period suggested. A team reporting +4.3% would have overstated the return by more than double, and nothing inside the test period could have revealed it.

Common misconceptions

The pre-period just needs to be long enough to match the markets.
It also establishes the baseline relationship the whole estimate rests on, and it has to span a full seasonal cycle for that relationship to be trustworthy. Markets that track each other over four quiet weeks may diverge completely at Christmas, and a matching built on the quiet weeks will not hold when it matters.
The cooldown is optional if you just want the campaign's effect.
Without it you cannot distinguish demand created from demand pulled forward, and those support opposite decisions. A campaign that borrows from next month shows a strong test-period lift and adds nothing over the quarter. The cooldown is the only phase in which that difference is visible.
A longer test period always gives a better estimate.
It gives a more precise estimate of a longer campaign, at proportionally more cost, and it does not fix the initial dilution from conversion lag — the right remedy there is to discard the lag window from the analysis. Extending the test also delays the cooldown and pushes back the next experiment.

Frequently asked questions

How long should the pre-period be?
At least one full seasonal cycle, which for most businesses means six months to a year. It has two jobs — matching the markets and establishing the baseline relationship the estimate depends on — and both degrade badly on a short window. Markets that track each other over a quiet month can diverge sharply during a peak, and a matching built on the quiet month will fail exactly when the test runs.
How long should the cooldown run?
Two to three times the adstock half-life, which for most consumer categories means two to four weeks. The purpose is to see whether treated markets fall below control once spend stops, which is the signature of pulled-forward demand. It also serves as the washout before the next test on the same markets, since residual advertising effects would otherwise contaminate it.
How do I handle the conversion lag at the start of a test?
Discard an initial window from the analysis rather than extending the test. If purchases typically follow exposure by two weeks, the first two weeks of the test period are largely recording conversions driven by pre-test activity, which dilutes the effect. Measure the lag from historical data, drop that window, and report the decision — it should be made before the test rather than after seeing which cut helps.

Related terms

  • Counterfactual forecast

    The dotted line on every geo chart — a prediction, not an observation, and the whole result rests on it.

  • Designated Market Area

    The 210 US television markets — the standard geo unit because the media buy already respects those boundaries.

  • Geo experiment

    Randomise regions instead of users — the way to test marketing that cannot be hidden from a person.

  • Scale-up vs scale-down test

    Add budget or switch it off — the direction decides which question you get an answer to.

Calculate it

  • Paired t-test

    Before-and-after or matched pairs — size the study on the difference SD, then test it.

  • Correlation test

    Pearson r or Spearman rho, with the Fisher-z interval that says how little a small sample knows.

Knowing the term is the easy part

Applying it to a live measurement problem is the part that goes wrong. If you are designing an experiment, reading a result you do not trust, or trying to work out what your marketing actually caused, that is the work we do.

References

  • Vaver, J., & Koehler, J. (2011). Measuring Ad Effectiveness Using Geo Experiments. Google Inc.
  • Kerman, J., Wang, P., & Vaver, J. (2017). Estimating Ad Effectiveness using Geo Experiments in a Time-Based Regression Framework. Google Inc.