Bell Statistics

How is an MMM calibrated?

MMM calibration uses experimental results to anchor a marketing mix model's estimates. Because an MMM fits correlations in historical spend, it can produce plausible and wrong coefficients — and a geo experiment supplies the ground truth that pins them down.

Also called
attribution model calibration, geo-calibrated MMM, experiment-informed priors, model calibration
Allon Korem

Written by Allon Korem

Chief Executive Officer

Last updated

In plain English

A marketing mix model estimates each channel's contribution from historical variation in spend and outcomes. That is fundamentally a correlational exercise, and it inherits the usual weakness: when two channels moved together, or a channel's budget barely varied, the data cannot separate their effects and the model produces a coefficient anyway. Calibration is the practice of anchoring those estimates to experiments that measured a channel's effect directly.

The problem it fixes is specific. An MMM will report a contribution for every channel in it, and the reported precision reflects how well the model fits rather than how well the effect is identified. Channels whose spend has been flat for two years, or that always move in step with another channel, get coefficients that are essentially assumptions dressed as estimates. Nothing in the output distinguishes those from the well-identified ones.

A geo experiment measures a channel's incrementality causally, which is precisely what the model cannot establish alone. In a Bayesian MMM the natural way to use that is as an informative prior on the relevant coefficient — the experiment says this channel returns around 2.8 with a certain uncertainty, and the model fits everything else subject to that constraint. The experiment pins one part of the model and the model interpolates the rest.

The effect is usually larger than expected because the coefficients are not independent. Constraining one channel to its measured value changes what the model can attribute to the others, since the total is bounded by observed outcomes. Calibrating a single overstated channel typically redistributes contribution across several, which is why one well-chosen experiment can improve an entire model rather than only one line of it.

The practical programme is to calibrate the channels where identification is weakest and the spend is largest, refresh annually because response curves shift, and treat an uncalibrated MMM's channel-level numbers as provisional. This is also what makes the combination of methods coherent rather than a portfolio of competing answers: experiments provide precise causal estimates on a few channels at a time, and the MMM extends that to everything, including channels no experiment can hold out.

The formula

The model, the constraint an experiment imposes on it, and why the effect propagates beyond the calibrated channel.

The model
y = Σ βᵢ · f( spendᵢ ) + trend + seasonality + ε

f carries adstock and diminishing returns. Every βᵢ is estimated from historical variation.

The identification problem
flat or collinear spend → β is barely constrained by data

The model reports a coefficient regardless, and the fit statistics do not flag it.

The calibration
prior on βᵢ centred on the experimental estimate, width from its interval

The experiment's uncertainty carries through rather than being treated as exact.

Why it propagates
coefficients are correlated; total contribution is bounded

Constraining one channel redistributes across others — see the correlation calculator.

Worked example

An MMM covering five channels is fitted, then recalibrated using two geo experiments — one on branded search, one on television. Neither channel's spend had varied much historically.

Uncalibrated MMM: branded search contribution
18% of modelled revenue
Geo test: branded search iROAS
0.9, implying ≈ 4% contribution
Uncalibrated MMM: television contribution
6%
Geo test: television iROAS
2.8, implying ≈ 15% contribution
Post-calibration: social contribution
moved from 11% to 14%
Post-calibration: model fit (R²)
0.91 → 0.89

Two experiments moved branded search from 18% to 4% and television from 6% to 15%, and redistributed contribution across channels that were never tested.

The two calibrated channels were the ones the model could least identify — branded search spend had tracked overall demand for years, and television budget had barely moved, so both coefficients were largely assumption. The model had attributed to branded search a great deal of demand that branded search was capturing rather than creating, which is the characteristic error when a channel's spend correlates with demand it does not cause. The row worth noticing is social moving from 11% to 14% without any experiment touching it: constraining two channels changes what the model can attribute elsewhere, so one calibration improves the whole allocation. The fit statistic falling slightly is expected and reassuring rather than a problem — the uncalibrated model fitted the history better precisely because it was free to assign contribution wherever the correlations pointed, and a slightly worse fit against a causally anchored structure is the trade being made deliberately.

Common misconceptions

An MMM with a good fit produces reliable channel estimates.
Fit measures how well the model reproduces historical outcomes, not how well individual channels are identified. A channel whose spend never varied gets a coefficient that the data barely constrains, and the fit statistic looks the same either way. Identification and fit are different properties and only one of them is reported.
Calibration means overriding the model with the experiment's number.
It constrains rather than replaces. The experiment enters as an informative prior carrying its own uncertainty, and the model fits everything else subject to it. Treating an experimental estimate as exact ignores that geo tests have wide intervals, and hard-coding it would discard information the model does have.
If the calibrated model fits worse, calibration made it worse.
A slightly worse fit is expected and usually a good sign. The uncalibrated model fitted history better because it was free to attribute contribution wherever correlations pointed, including to channels that caused nothing. Trading a little fit for causal anchoring is the point of the exercise.

Frequently asked questions

Which channels should I run experiments on for calibration?
Those where spend is large and identification is weakest — channels whose budget has been flat, or that always move alongside another. Those are where the model is least constrained by data and most likely to be wrong. Branded search and television are frequent candidates: both often have stable spend and both are commonly mis-estimated in opposite directions.
How often does an MMM need recalibrating?
Annually as a baseline, and sooner after a material change to strategy, creative or the competitive landscape. Response curves shift as audiences saturate and competitors move, so an experiment from two years ago describes a channel that may no longer behave the same way. Running one or two geo tests a year on a rotating set of channels keeps the model anchored without an unmanageable measurement programme.
How should the experiment's uncertainty be handled?
Carry it through as the width of the prior. Geo tests produce wide intervals, and a prior centred on the point estimate with a narrow width overstates what the experiment established — which replaces one overconfident number with another. A prior whose spread matches the experiment's interval lets the model tighten the estimate where the historical data genuinely helps.

Related terms

  • Attribution window

    A dial that changes every channel's reported performance — and platforms do not set it the same way.

  • Last-click attribution

    All the credit to the last touch — reproducible, universally understood, and wrong in a predictable direction.

  • Marketing mix modelling

    One regression across every channel, built on aggregate data — no tracking, and strong assumptions.

  • Multi-touch attribution

    Credit spread across the path — better than last-click, and still describing correlation rather than cause.

Calculate it

  • Correlation test

    Pearson r or Spearman rho, with the Fisher-z interval that says how little a small sample knows.

  • Paired t-test

    Before-and-after or matched pairs — size the study on the difference SD, then test it.

Knowing the term is the easy part

Applying it to a live measurement problem is the part that goes wrong. If you are designing an experiment, reading a result you do not trust, or trying to work out what your marketing actually caused, that is the work we do.

References