Bell Statistics

What is a halo effect?

A halo effect is a benefit that lands outside the thing you changed — a promoted product lifting sales of others, or a campaign for one line raising the whole brand. A metric scoped to the change alone misses it and understates the result.

Also called
spillover benefit, cross-category lift, indirect effect
Allon Korem

Written by Allon Korem

Chief Executive Officer

Last updated

In plain English

Advertise one product and customers buy others. Improve one page and satisfaction rises across the journey. Run a brand campaign and a competitor's paid search gets more expensive because more people are typing your name. In each case the benefit is real and lands somewhere other than the thing that was changed — so a metric drawn tightly around the change reports less than actually happened.

This is cannibalization with the sign reversed, and the two share one root cause: a metric scoped more narrowly than the effect. Cannibalization means the narrow metric flatters the result because gains were taken from elsewhere. A halo means it understates the result because gains landed elsewhere. Both are failures of scope rather than of statistics, and both are found the same way — by measuring at a level wide enough to contain the spillover.

The consequences fall unevenly, and predictably. Changes whose effects are concentrated and immediate measure well: a checkout fix shows up in checkout conversion. Changes whose effects are diffuse measure badly: brand advertising, content, trust signals and quality improvements all tend to produce small distributed gains across many surfaces, none of which is individually significant. An organisation that funds only what measures cleanly will systematically over-invest in the first kind and under-invest in the second, without anyone deciding to.

In media measurement this is the strongest argument against channel-level attribution. Upper-funnel advertising works partly by making later touchpoints more effective — people who saw the brand campaign convert better on search — and a model that assigns credit to the last click books that entire benefit to search. The brand campaign then looks like a poor performer and gets cut, after which search performance quietly declines. Only a method that sees all channels at once catches this: a geo experiment that withholds the campaign from some markets, or marketing mix modelling fitted across the whole portfolio.

The practical difficulty is that widening the metric costs power. Total revenue is far noisier than one product's revenue, so an experiment honest enough to capture the halo may be unable to resolve it. That is a real constraint rather than an argument against trying: the usual resolution is to measure the narrow metric for the decision and the wide one for the programme, checking periodically whether the sum of individually measured effects matches the aggregate.

The formula

The same decomposition as cannibalization, with the second term positive. What differs is where you have to look to find it.

The decomposition
Δ total = Δ focus + Δ elsewhere

Cannibalization is Δ elsewhere < 0; a halo is Δ elsewhere > 0. The narrow metric sees only the first term either way.

Halo ratio
( Δ total − Δ focus ) / Δ focus

How much extra arrives for each unit measured directly. Ratios above 0.5 are common for upper-funnel media.

Why it is hard to detect
small effect spread across k surfaces, each with its own σ

k small non-significant gains can sum to a large real one that no individual test can resolve.

What captures it
measure at market level, not channel level

A holdout across whole geographies contains every spillover — see the ANOVA calculator for multi-market comparisons.

Worked example

A retailer runs a four-week brand campaign for its outerwear line. Last-click attribution is used alongside a geo experiment in which the campaign is withheld from 18 matched markets. The question is what the campaign was worth, and the two methods disagree substantially.

Last-click revenue attributed to the campaign
£410,000
Campaign spend
£520,000
Last-click ROAS
0.79 — apparently loss-making
Geo test: outerwear lift in treated markets
£498,000
Geo test: other categories lift
£372,000
Geo test: total incremental revenue
£870,000

Last click values the campaign at £410,000 against £520,000 of spend. The geo test measures £870,000 of incremental revenue, of which 43% landed outside the advertised category.

On the attribution number this campaign gets cut. On the geo number it returns £1.67 for every pound and should be expanded. The gap is almost entirely halo: £372,000 of the lift appeared in categories the campaign never mentioned, which last-click has no mechanism for seeing because those purchases arrived through search and direct visits with no campaign touchpoint. Two cautions before treating the geo number as settled. The 18-market design has its own uncertainty and the interval around £870,000 is wide — this is one measurement, not a constant. And a four-week campaign may pull forward purchases that would have happened later, which a four-week measurement window would count as incremental; extending the post-period is what distinguishes genuine growth from timing. What the comparison does establish firmly is that the attribution figure is a floor rather than an estimate.

Common misconceptions

If the effect were real, it would show up in the metric we measured.
Only if the metric is wide enough to contain it. A diffuse benefit spread across a dozen surfaces produces a dozen small movements, none individually significant, that sum to something substantial. The narrow metric is not measuring a smaller effect — it is measuring a fraction of the effect.
Multi-touch attribution captures halo effects because it credits multiple touchpoints.
It distributes credit among touchpoints that appear in observed conversion paths, which is a different problem. A halo works partly by making later touchpoints more effective and partly by generating conversions with no upper-funnel touch recorded at all. Neither is visible in the path data, so no allocation rule over that data can recover it.
Halo effects mean you should always measure the widest possible metric.
The widest metric is usually too noisy to resolve anything, so insisting on it means never detecting anything. The workable structure is to decide individual experiments on a sensitive narrow metric and periodically check the aggregate — if the sum of measured wins does not appear in the totals, the scope is wrong somewhere.

Frequently asked questions

How do I measure a halo effect?
With a design whose unit is wide enough to contain the spillover, which usually means a geo experiment: withhold the change from whole markets and compare total outcomes rather than category outcomes. Within a product, the equivalent is randomising at a level above where the spillover happens — by account rather than by user if the effect crosses between colleagues. What will not work is any channel-level or surface-level metric, because the effect by definition leaves that scope.
Why does brand advertising measure so badly on attribution?
Because most of its value arrives indirectly. It makes later touchpoints more effective and generates conversions where no brand ad appears in the recorded path, and attribution can only allocate credit among touchpoints it observed. The predictable result is that upper-funnel spend looks unprofitable, gets cut, and lower-funnel performance declines a quarter later — a pattern common enough to be worth watching for.
How do I know if my programme is missing halo effects?
Compare the sum of your measured wins against the aggregate metrics over the same period. If a year of experiments claims 20% of cumulative improvement and the company metrics show 5%, something is wrong with the scope — though that gap can equally come from novelty effects or optimistic proxies, so it identifies a problem rather than its cause. A long-term holdout is what turns that comparison into a measurement.

Related terms

  • Cannibalization

    Moving demand and calling it growth — the failure that only a total-level metric can see.

  • Hangover effect

    The cost of relearning, mistaken for a worse product — and the reason a good change can lose its first week.

  • Incrementality

    The conversions that would not have happened anyway — and the gap between that and what platforms report.

  • Novelty effect

    Curiosity, measured and mistaken for improvement — and the reason a strong week-one result is the least trustworthy kind.

Calculate it

  • One-way ANOVA

    Three or more independent groups on one continuous outcome — size it, then run the F test.

  • Correlation test

    Pearson r or Spearman rho, with the Fisher-z interval that says how little a small sample knows.

Knowing the term is the easy part

Applying it to a live measurement problem is the part that goes wrong. If you are designing an experiment, reading a result you do not trust, or trying to work out what your marketing actually caused, that is the work we do.