Bell Statistics

What is a conversion lift study?

A conversion lift study is a randomised experiment run inside an advertising platform, holding out a portion of the eligible audience from seeing the ads. It measures incrementality properly, within the boundary of that one platform.

Also called
lift study, platform lift test, conversion lift test, brand lift study
Allon Korem

Written by Allon Korem

Chief Executive Officer

Last updated

In plain English

A conversion lift study is a genuine randomised experiment: the platform splits the eligible audience, shows ads to one group and withholds them from the other, and compares conversion rates. Because assignment is random, the difference is causal — this is a real measurement of incrementality rather than an attribution model, and it deserves more credit than it usually gets from people sceptical of platform-reported numbers.

Its advantages over a geo experiment are practical. It randomises individuals rather than markets, so the sample is enormous and the estimate is far more precise. It requires no media reorganisation. And the holdout is managed by the platform, so the mechanics are handled. For measuring one platform's contribution it is often the most precise tool available.

The limitations are all about the boundary. It measures the platform's incrementality relative to everything else you are doing, which means the result is conditional on the rest of your media mix. It cannot see cross-channel effects — the halo from a campaign there landing in search or direct, or the cannibalization of one channel by another. And it necessarily excludes anyone the platform could not have reached, so the population is the platform's addressable audience rather than your customer base.

Then there is the conflict of interest, which is real and should be handled proportionately rather than used to dismiss the method. The platform designs the study, defines eligibility, runs the randomisation and computes the result. That is not evidence of manipulation and it is a reason to check what is checkable: confirm the holdout size, ask for the confidence interval rather than accepting a point estimate, and where the numbers matter enough, validate against an independent geo test on the same channel.

The sensible position is that lift studies and geo tests are complements. Use a lift study for precise within-platform measurement and iteration; use a geo test when you need cross-channel effects, an independent read, or a channel the platform cannot hold out. Where both are run and agree, that agreement is worth more than either alone — and where they disagree, the geo test is generally measuring the wider quantity.

The formula

The estimator is an ordinary randomised comparison. What needs attention is the population it applies to.

The lift
lift = ( conv_exposed − conv_holdout ) / conv_holdout

A randomised difference, so causal within the study population.

The population
users the platform deemed eligible and could have reached

Not your customer base. The result generalises only to that audience.

What it cannot see
effects landing outside the platform

Cross-channel halo and cannibalisation are outside the boundary by construction.

Converting to iROAS
iROAS = incremental revenue / spend in the study

Requires revenue per conversion — see the one-proportion z-test calculator.

Worked example

An advertiser runs a platform conversion lift study on a social channel and, in the same quarter, a geo test that switches the same channel off in matched markets. The two results are compared.

Lift study: holdout size
10% of eligible audience
Lift study: measured lift
+14.2%, 95% CI +12.1% to +16.3%
Lift study: implied iROAS
3.9
Geo test: measured lift on total conversions
+7.8%, 95% CI +2.4% to +13.2%
Geo test: implied iROAS
2.2
Search conversions in geo test markets
−2.9% when social was off

The lift study reports 3.9 iROAS with a tight interval; the geo test reports 2.2 with a wide one. Both are correct about different things.

The last row reconciles them. When social was switched off in the geo markets, search conversions fell by 2.9% — some of the conversions the lift study counts as social's own were being handed to search in the wider picture, and switching social off removed them from both places. The lift study is precise and boundary-limited: it correctly measures social's incrementality relative to everything else running. The geo test is imprecise and complete: it measures what happens to total conversions, cross-channel effects included. For a decision about whether to keep the channel, the geo figure of 2.2 is the relevant one. For iterating on social creative and targeting week to week, the lift study's precision is far more useful. The intervals also tell a story worth respecting — the geo test's range of 2.4% to 13.2% is honest about how little a market-level design can pin down.

Common misconceptions

Platform lift studies are just attribution with better marketing.
They are randomised experiments with a genuine holdout, which makes them causal evidence rather than an allocation rule. That is a real methodological difference from attribution and it deserves recognition. The legitimate criticisms concern the boundary of what they measure and who runs them, not whether the design is sound.
A lift study measures the channel's total contribution.
It measures the channel's contribution within its own boundary, conditional on everything else you are running. Cross-channel effects — demand shifted to or from search and direct — are invisible to it by construction. That is precisely the quantity a geo test captures and a lift study cannot.
Because the platform runs it, the result cannot be trusted.
The conflict is real and the design is sound, so the proportionate response is verification rather than dismissal. Check the holdout size, insist on confidence intervals, and where the stakes justify it validate against an independent geo test. Discarding the most precise measurement available is not a conservative choice.

Frequently asked questions

Should I run a lift study or a geo test?
A lift study for precise measurement within one platform and for iterating quickly, since it randomises individuals and gives tight intervals. A geo test when you need cross-channel effects included, an independent read, or a channel the platform cannot hold out — television, for instance. They are complements, and where both run and agree the conclusion is much stronger than either alone.
What should I check on a platform lift study?
The holdout size and how eligibility was defined, since that determines who the result generalises to. The confidence interval rather than the point estimate. Whether the conversion window matches how you measure elsewhere. And whether the reported lift is relative or absolute, because a relative figure on a low base rate can sound far more impressive than the absolute movement warrants.
Why can't a lift study see cross-channel effects?
Because both its arms sit inside your wider media mix. Holdout users still see your search ads, emails and everything else, so the study measures the platform's contribution on top of all that. If switching the channel off would shift demand to search, that shift affects both arms equally and cancels. Only an experiment whose unit contains all the channels — a geo test — can observe it.

Related terms

  • Geo experiment

    Randomise regions instead of users — the way to test marketing that cannot be hidden from a person.

  • Ghost ads

    Log the ad you would have shown instead of showing it — exposure-matched control, and no wasted spend.

  • Incremental CPA

    Cost per conversion you actually caused — the number bids should be set against, and rarely are.

  • iROAS

    Return on spend counting only what the advertising caused — routinely a fraction of the platform's number.

Calculate it

  • One-proportion z-test

    Test one observed rate against a fixed target — an SLA, a benchmark, a contractual floor.

  • A/B test sample size

    Size a two-proportion experiment before you launch, then read the lift and its interval once it lands.

Knowing the term is the easy part

Applying it to a live measurement problem is the part that goes wrong. If you are designing an experiment, reading a result you do not trust, or trying to work out what your marketing actually caused, that is the work we do.