Bell Statistics

Leading or lagging indicator?

A lagging indicator reports an outcome after it has settled, such as quarterly churn. A leading indicator moves earlier and is believed to predict it. The pair exists because the measures that matter most are the ones that arrive too late to act on.

Also called
leading indicator, lagging indicator, leading vs lagging metrics
Allon Korem

Written by Allon Korem

Chief Executive Officer

Last updated

In plain English

The distinction is about timing relative to a decision. A lagging indicator measures an outcome once it has happened — quarterly revenue, annual churn, six-month retention. It is authoritative and it arrives too late to change anything. A leading indicator moves earlier and is believed to anticipate that outcome — activation rate, week-one engagement, trial-to-paid conversion — so it can inform a decision while the decision is still open.

The trade is always the same. Lagging indicators are what you actually care about and cannot steer by; leading indicators are actionable and their connection to what you care about is an assumption. That assumption is the whole substance of the pair, and it is exactly the claim a proxy metric makes — the two concepts overlap heavily, with the leading/lagging framing emphasising timing and the proxy framing emphasising substitution.

What makes a leading indicator genuinely leading rather than merely early is that it predicts the outcome under intervention, not just in historical correlation. Support ticket volume correlates with churn, but that may be because unhappy customers do both rather than because tickets cause departures — in which case suppressing tickets by making support harder to reach would move the leading indicator and worsen the lagging one. Every leading indicator carries this risk, and it is the reason the pairing needs periodic checking rather than one-off validation.

In practice a healthy measurement system runs both and uses them for different jobs. Leading indicators decide experiments and weekly operating reviews, because they move fast enough to be informative on that cadence. Lagging indicators validate the leading ones and settle the question of whether the programme is actually working — which requires deliberately holding some slow readout, whether that is a long-term holdout or simply revisiting shipped changes after a quarter.

The failure this structure prevents is a specific and common one: a year of experiments all reporting wins on leading indicators, with the lagging numbers flat. Every individual test was correctly run and correctly analysed. What went wrong is upstream of any of them — the assumed link never held, and because nothing was measuring the lagging side, nothing could say so. The only defence is measuring both and comparing them on purpose.

The formula

No formula defines the pair, but the quantity that decides whether a leading indicator is worth using can be written down, and it is not the correlation people usually quote.

What is usually cited
ρ( leading_t , lagging_t+k )

Correlation at a lag, across users or periods. Necessary and nowhere near sufficient — it cannot distinguish prediction from a shared cause.

What actually matters
Δ lagging / Δ leading, across past interventions

The transfer rate under intervention. Requires having shipped changes and looked back at the slow outcome.

The failure mode
leading ↑, lagging unchanged

The signature of a shared cause rather than a causal link — see confounding.

Why the fast one is used
n ∝ σ² / Δ²

A lagging indicator is slow AND noisy, so it usually cannot power an experiment at all — see the sample size calculator.

Worked example

A B2B software company treats weekly active seats as its leading indicator for annual renewal, the lagging one. Renewal is 84% and can only be observed once a year. Over eighteen months they run experiments on seat activation and want to know whether the leading indicator is earning its authority.

Lagging: annual renewal rate
84%, observable once per account per year
Leading: weekly active seats
readable in 14 days
Account-level correlation
0.63
Experiments that lifted active seats
7, mean +6.8%
Renewal change in those cohorts
+1.1 pp against a predicted +5.7 pp
Experiments where renewal did not move
3 of 7

The leading indicator is real but weak: roughly a fifth of the predicted renewal effect materialised, and three of seven changes transferred nothing.

A correlation of 0.63 looks like strong validation and is doing much less work than it appears to. The transfer rate of about 20% is the number that should drive forecasting, and it is only knowable because the company waited a year on seven cohorts and looked. The three non-transferring experiments are worth examining individually rather than averaging away — reviewing them, two had increased seat activation through admin bulk-invites, which adds seats that were never going to be used and moves the leading indicator without touching the thing it was supposed to predict. That is the shared-cause failure in its ordinary clothing. The practical response is not to abandon the leading indicator, which is still the only thing fast enough to steer by, but to narrow it: active seats that logged in twice or more, which is harder to manufacture and should transfer better. Then wait another year and check again.

Common misconceptions

A leading indicator is just an early version of the lagging one.
It is a different quantity that is believed to predict the lagging one, and the belief can be wrong in a specific way — both may be driven by a third factor rather than one causing the other. When that is the case, interventions move the leading indicator and leave the lagging one exactly where it was.
A strong correlation between the two proves the leading indicator works.
Correlation across users or periods is consistent with prediction and equally consistent with a shared cause, and only the first survives intervention. The evidence that counts comes from changes you actually shipped: did the ones that moved the leading indicator also move the lagging one, and by how much.
Once you have good leading indicators you can stop measuring the lagging ones.
Then nothing can tell you when the link breaks, and it does break as the product and user base change. The failure is silent by construction — every experiment keeps reporting wins on the fast metric. Keeping a slow readout is what converts that from an invisible drift into a detectable one.

Frequently asked questions

Which should I use to decide an experiment?
The leading indicator, almost always, because a lagging one cannot be read inside an experiment's duration and is usually too noisy to power a test on even if you waited. What the lagging indicator is for is validating that choice — periodically checking that the changes which moved the leading metric also moved the outcome. Use the fast one to steer and the slow one to confirm you are steering somewhere real.
Is a leading indicator the same as a proxy metric?
They overlap almost entirely and emphasise different aspects. Leading and lagging is about timing — one arrives before the decision, the other after. Proxy and outcome is about substitution — one stands in for the other because the real thing cannot be measured in time. In practice a leading indicator used to decide experiments is functioning as a proxy, and it inherits every validation requirement that comes with one.
How often should I check that a leading indicator still predicts?
At least annually, and whenever something structural changes — a new acquisition channel, a repositioning, a major product shift. The check needs shipped changes and a slow readout, so it has to be planned rather than performed on demand: hold a long-term measurement on a sample of releases so the comparison is available when you want it. Retrofitting it after suspicion arises usually means the data was never collected.

Related terms

  • Incrementality

    The conversions that would not have happened anyway — and the gap between that and what platforms report.

  • North star metric

    A coordination tool for the company, not a decision rule for an experiment — and confusing the two is the usual mistake.

  • Primary metric

    The one number the decision hangs on — nominated before the data arrives, which is the entire point.

  • Proxy metric

    A stand-in for the outcome you cannot wait for — and the correlation it rests on is an assumption, not a finding.

Calculate it

  • Correlation test

    Pearson r or Spearman rho, with the Fisher-z interval that says how little a small sample knows.

Knowing the term is the easy part

Applying it to a live measurement problem is the part that goes wrong. If you are designing an experiment, reading a result you do not trust, or trying to work out what your marketing actually caused, that is the work we do.

References