Bell Statistics

What is cannibalization?

Cannibalization is when a gain in one place is taken from another rather than created. The measured metric improves, the total does not move, and the experiment reports a win because it was only ever looking at one side of the ledger.

Also called
cannibalisation, substitution effect, demand shifting, channel cannibalization
Allon Korem

Written by Allon Korem

Chief Executive Officer

Last updated

In plain English

Promote a product on the homepage and its sales rise. Whether that is growth depends entirely on where the sales came from. If customers who would have bought something else bought this instead, the product's metric improved and the business gained nothing — demand moved rather than increased. That is cannibalization, and it is invisible to any measurement scoped to the thing that improved.

It appears in three recognisable places. In product experiments, a promoted item or surface wins at the expense of others: a bigger recommendation module lifts recommendation clicks and takes them from search. In media, a channel takes credit for conversions another channel would have delivered — branded search is the classic case, capturing purchases from people already intent on buying, which is why its measured ROAS is spectacular and its incrementality often close to zero. In pricing and merchandising, a discounted tier pulls customers down from a higher-margin one.

What makes it dangerous is that it is not a measurement error. Every individual number is correct. The promoted product really did sell more; the channel really did drive those conversions in the sense that its ad was the last thing touched. The failure is in scope — the metric was defined narrowly enough that substitution falls outside it, so the experiment is answering a question nobody meant to ask.

The fix is structural rather than statistical: measure at the level where substitution is contained. If demand can move between products, the metric is total revenue rather than the promoted product's revenue. If it can move between channels, no channel-level measurement will do and the answer requires a geo experiment or a model that sees all channels at once. The general rule is that the metric has to be at least as wide as the substitution you are worried about, and widening it costs sensitivity — total revenue is far noisier than one product's revenue, which is exactly why teams are tempted to measure narrowly.

Its opposite is the halo effect, where a change lifts things outside its own scope rather than draining them. Both are the same measurement failure — a metric scoped more narrowly than the effect — and both are reasons that a portfolio of individually successful experiments can add up to a flat quarter. When that pattern appears, cannibalization is the first thing to look for.

The formula

One subtraction, applied at the widest level demand can move within. Everything difficult about cannibalization is choosing that level rather than doing the arithmetic.

The rate
cannibalization rate = ( gain_focus − gain_total ) / gain_focus

0 means everything was incremental; 1 means the entire gain came from elsewhere.

What the narrow metric reports
Δ focus = incremental + substituted

The two components are indistinguishable from inside the narrow metric. Only the total separates them.

The measurement rule
metric scope ⊇ substitution scope

If demand can move between products, measure total revenue. Between channels, measure at market level — see geo experiment.

What widening costs
n ∝ σ²_total / Δ²

The total is much noisier than the part, so the honest metric needs more traffic — see the ANOVA calculator for multi-arm cases.

Worked example

A retailer tests a homepage banner promoting its own-brand range. The experiment is judged on own-brand revenue per visitor, and the team also has total revenue per visitor available as a guardrail. Both are measured across 210,000 visitors per arm.

Own-brand revenue per visitor
£3.10 → £3.98 (+28.4%)
Branded-supplier revenue per visitor
£8.40 → £7.71 (−8.2%)
Total revenue per visitor
£11.50 → £11.69 (+1.7%, p = 0.21)
Cannibalization rate
(0.88 − 0.19) / 0.88 = 78%
Own-brand gross margin
42%
Branded-supplier gross margin
19%

78% of the own-brand gain came out of branded-supplier sales. Total revenue moved 1.7% and is not statistically distinguishable from zero.

Judged on its stated metric this is a 28% win and one of the year's best results. Judged on total revenue it is roughly nothing. The interesting wrinkle is the margin columns: shifting £0.69 of revenue from 19% margin to 42% margin adds about £0.16 of gross profit per visitor even with no revenue growth at all, so the change may well be worth shipping — on a completely different rationale from the one the experiment was set up to test. That is the useful lesson. Cannibalization does not automatically mean a change is bad; it means the metric was too narrow to evaluate it, and the real decision needs a wider frame. What would have been indefensible is reporting +28.4% as incremental growth. Note the cost of honesty here: total revenue is noisy enough that even a genuine 1.7% lift is not resolvable at 210,000 visitors per arm.

Common misconceptions

The experiment measured a real lift, so the change created value.
It measured a real lift in the thing it looked at. Whether value was created depends on what happened outside that scope, and a narrow metric cannot distinguish demand that was created from demand that moved. Only a total-level comparison separates the two.
Cannibalization means the change should not ship.
Not necessarily — it means the stated justification was wrong. Shifting demand towards higher-margin products, or towards a channel that costs less to serve, can be clearly worthwhile even at complete substitution. What has to change is the metric the decision is made on, not automatically the decision.
Attribution data will show whether channels are cannibalising each other.
Attribution allocates credit for conversions that happened; it has no view of what would have happened otherwise. A channel that captures purchases already destined to occur looks excellent in every attribution model. Separating capture from creation needs an experiment that withholds the channel — a geo test or a conversion lift study.

Frequently asked questions

How do I detect cannibalization in an A/B test?
Measure the total alongside the part, and read them together. If the promoted item gains 28% while the category total moves 2%, substitution accounts for most of it. This only works if the total metric is in the experiment from the start — adding it after seeing a suspicious result means it was not powered for and probably cannot resolve the difference. Treating the total as a mandatory guardrail on any promotion test is the cheap version of this discipline.
Why is branded search the classic cannibalization example?
Because people searching for your brand name have already decided to visit you, so an ad shown to them captures a conversion that would very likely have happened through the organic link. Attribution credits the ad, ROAS looks outstanding, and the incremental contribution is often near zero. The only way to find out is to switch branded search off in some markets and compare — which is what geo experiments exist for.
How is cannibalization different from a halo effect?
They are the same measurement failure with opposite signs. Cannibalization means a narrow metric overstates the effect because gains were taken from elsewhere; a halo effect means it understates the effect because benefits landed outside the measured scope. Both come from a metric scoped more narrowly than the change's real reach, and both are found by measuring at the level where the substitution or spillover is contained.

Related terms

  • Halo effect

    The gains that land where nobody was measuring — cannibalization's mirror image, and the reason good work looks flat.

  • Hangover effect

    The cost of relearning, mistaken for a worse product — and the reason a good change can lose its first week.

  • Novelty effect

    Curiosity, measured and mistaken for improvement — and the reason a strong week-one result is the least trustworthy kind.

  • Primary metric

    The one number the decision hangs on — nominated before the data arrives, which is the entire point.

Calculate it

  • One-way ANOVA

    Three or more independent groups on one continuous outcome — size it, then run the F test.

  • Two-sample t-test

    Compare the average of two independent groups — plan the sample size, then test the result.

Knowing the term is the easy part

Applying it to a live measurement problem is the part that goes wrong. If you are designing an experiment, reading a result you do not trust, or trying to work out what your marketing actually caused, that is the work we do.