Bell Statistics

What is adstock?

Adstock is the modelled carry-over of advertising: the assumption that a campaign keeps working after it stops, with its influence decaying week by week. It converts a spend series into an accumulated-exposure series before that enters a model.

Notation
λ
Also called
carryover effect, advertising decay, lagged effect
Allon Korem

Written by Allon Korem

Chief Executive Officer

Last updated

In plain English

Advertising does not deposit its whole effect in the week it airs. Somebody sees a television spot on Tuesday and buys three weeks later; somebody else remembers the brand two months on. A model that regresses this week's sales on this week's spend assumes all of that away, and it will systematically understate every channel with a long tail — which is to say every brand-building channel. Adstock is the standard repair: transform the spend series into a series of accumulated, decaying exposure, and use that as the model's input.

The usual form is geometric. This week's adstock is this week's spend plus a fraction λ of last week's adstock, so the effect of a single burst decays by a constant proportion each period. λ near zero means the channel acts almost entirely on impact; λ near one means it lingers for months. Typical estimates run around 0.1-0.3 for paid search, where intent is immediate, and 0.5-0.8 for television, where the effect accumulates and persists — but these are starting points, not constants, and they differ by category, creative and audience.

It matters because getting λ wrong redistributes credit between channels. Set television's decay too low and its measured contribution collapses, because most of what it did falls outside the window the model is looking at, and whatever channel happens to run later — usually search — absorbs the credit. Set it too high and television appears to be responsible for sales it had nothing to do with. Since search often runs continuously while television is flighted, this single parameter can reverse the apparent ranking of the two largest lines in a media plan.

Two refinements matter in practice. Advertising rarely peaks in the same week it runs — there is a lag before the effect builds, particularly for anything requiring consideration — so a delayed adstock that peaks at week one or two often fits better than pure geometric decay. And a Weibull specification allows the decay shape itself to vary rather than fixing it as constant-proportional, which suits campaigns whose effect builds and then falls away rather than decaying monotonically from the start.

Ideally λ is estimated rather than assumed, and in a Bayesian marketing mix model it is given an informative prior and inferred alongside everything else. In practice it is often weakly identified, because separating carry-over from diminishing returns and from ordinary seasonality demands more variation than most spend histories contain. When it cannot be identified, say so and run sensitivity analysis across a plausible range — a conclusion that holds only at λ = 0.7 is a conclusion about the assumption, not about the channel.

The formula

One recursion and its consequences. Everything about how long a campaign keeps working follows from λ.

Geometric adstock
A_t = x_t + λ · A_{t−1}, 0 ≤ λ < 1

x_t is spend or impressions in period t. The recursion is what turns a burst into a decaying tail.

Half-life
half-life = ln(0.5) / ln(λ) periods

λ = 0.5 gives a one-week half-life; λ = 0.8 gives about three weeks; λ = 0.9 about seven. The intuitive way to sanity-check a fitted value.

Total multiplier
Σ λ^k = 1 / (1 − λ)

λ = 0.8 means a burst eventually delivers five times its immediate exposure. Normalise by this factor if you want adstock on the same scale as spend.

Delayed adstock
A_t = Σ_k w_k · x_{t−k}, w_k = λ^{(k − θ)²}

θ is the peak lag. Fits channels whose effect builds for a week or two before decaying, which is most brand advertising.

Worked example

A brand runs a single £500,000 television burst in week 10 and nothing before or after. An analyst compares what a model sees under three assumptions about decay, holding everything else fixed.

Spend, week 10
£500,000
Spend, all other weeks
£0
λ = 0.0, exposure in weeks 10-14
500k, 0, 0, 0, 0
λ = 0.5, exposure in weeks 10-14
500k, 250k, 125k, 63k, 31k
λ = 0.8, exposure in weeks 10-14
500k, 400k, 320k, 256k, 205k
Observed sales lift
spread fairly evenly across weeks 10-14

At λ = 0, the model can only credit week 10, and the lift in weeks 11-14 is attributed to whatever else was running. At λ = 0.8 it credits the burst across all five weeks.

The sales data are identical in all three cases; only the assumption differs, and it changes which channel gets paid. With λ = 0 the model sees television active for one week and sales elevated for five, so four weeks of lift get absorbed by search and by the intercept — search was running continuously, so it will happily take the credit, and the resulting plan shifts budget from television to search on the strength of an assumption rather than an observation. The honest procedure is to let the model estimate λ from a history containing several flights, and where it cannot, to report the channel's return across a range of λ so the reader can see how much of the answer is the data and how much is the prior.

Common misconceptions

Adstock is a fixed industry constant per channel.
It varies by category, creative, audience and purchase cycle. A car brand and a snack brand should not share a television decay rate, because one purchase is considered over months and the other over seconds. Published ranges are useful priors and a poor substitute for estimating it from your own history.
Applying adstock inflates a channel's measured contribution.
It reallocates rather than inflates. Total sales are fixed, so credit shifted towards a lingering channel comes from somewhere else — usually from whatever runs continuously and happens to coincide with the tail. The right value is whatever the data support, not whichever direction moves the number you prefer.
We can just add a few lagged spend terms instead.
You can, and lagged spend variables are almost perfectly correlated with each other, so the coefficients become unstable and can flip sign. Adstock imposes a shape on the decay and estimates one parameter instead of many, which is the whole reason the transformation exists rather than a general distributed lag.

Frequently asked questions

What are typical adstock decay rates by channel?
As rough priors: paid search around 0.1-0.3, since the intent is immediate; display and paid social around 0.3-0.5; television, radio and out-of-home around 0.5-0.8. Longer purchase cycles push all of these up, so a considered purchase can justify television values above 0.8. Treat published figures as starting points for a prior and estimate from your own data wherever the spend history contains enough flighting to identify it.
How do I estimate adstock rather than assume it?
Fit it as a parameter, ideally within a Bayesian model that gives it an informative prior and infers it jointly with the response curves. A cruder alternative is a grid search across candidate values, choosing by out-of-sample fit rather than in-sample. Either way, identification requires that spend actually varied — a channel that ran at a constant level for two years contains almost no information about how long its effect persists.
What is the difference between adstock and saturation?
Adstock is about time: how long the effect of spend persists after it runs. Saturation is about level: how the response flattens as spend within a period rises. They are applied in sequence — accumulate exposure over time first, then pass that through the response curve — and they are frequently confused because both make the relationship between spend and sales non-linear.

Related terms

  • Diminishing returns

    The tenth million does less than the first — and why average ROAS is the wrong number to budget on.

  • Marketing mix modelling

    One regression across every channel, built on aggregate data — no tracking, and strong assumptions.

  • Multicollinearity

    When predictors move together the model cannot separate them — good predictions, meaningless coefficients.

  • Overfitting

    A model that memorised the noise — excellent on the data it saw, useless on the data it will meet.

  • Regression analysis

    Fit a line through the data — and the phrase 'holding everything else fixed' is where the trouble starts.

Calculate it

  • Correlation test

    Pearson r or Spearman rho, with the Fisher-z interval that says how little a small sample knows.

  • Two-sample t-test

    Compare the average of two independent groups — plan the sample size, then test the result.

Knowing the term is the easy part

Applying it to a live measurement problem is the part that goes wrong. If you are designing an experiment, reading a result you do not trust, or trying to work out what your marketing actually caused, that is the work we do.

References

  • Broadbent, S. (1979). One Way TV Advertisements Work. Journal of the Market Research Society, 21(3), 139-166.
  • Jin, Y., Wang, Y., Sun, Y., Chan, D., & Koehler, J. (2017). Bayesian Methods for Media Mix Modeling with Carryover and Shape Effects. Google Inc. Technical Report.