Bell Statistics

What is a north star metric?

A north star metric is the single measure a company uses to align its teams on what progress means. It operates at organisational scale over quarters, which makes it a coordination device rather than something any individual experiment can be judged against.

Also called
north star, company metric, single key metric
Allon Korem

Written by Allon Korem

Chief Executive Officer

Last updated

In plain English

A north star metric is a company-level answer to "what does progress look like". Nights booked for a marketplace, weekly active teams for a collaboration product, hours watched for a streaming service. Its purpose is coordination: when six teams are choosing between roadmaps, a shared definition of progress stops them optimising six different things that pull against each other.

The distinction that matters most is between this and a primary metric. A primary metric decides one experiment — it must be sensitive enough to move detectably in two or three weeks, and it is nominated per test. A north star operates over quarters at company scale and is deliberately slow, broad and stable. It is almost always far too insensitive to serve as an experiment's decision rule: a checkout improvement that adds real value will not visibly shift weekly active teams, and a test powered to detect that would need more traffic than the company has.

The connection between them should be explicit rather than assumed. A well-constructed programme can draw a line from each experiment's primary metric up through intermediate outcomes to the north star, and can say roughly what a 1% move in the primary is worth at the top. Most organisations cannot draw that line, which is when the north star becomes decorative — quoted in all-hands presentations and connected to nothing anyone does on a Tuesday.

A good one has three properties. It reflects value delivered to customers rather than value extracted from them, so that improving it is genuinely good news. It is hard to move without doing real work, which rules out most vanity counts. And it is resistant to gaming, because any metric given organisational authority will be optimised whether or not anyone intends to cheat. "Hours watched" is the standard cautionary example: it can be improved by making content genuinely better and equally by making it harder to stop watching, and the metric cannot tell those apart.

The usual accompaniment is a small set of counter-metrics that must not degrade — the organisational equivalent of guardrail metrics. A north star for engagement paired with a watch on churn and support contacts is much harder to game than one on its own, because the cheap ways to move the headline number tend to show up immediately in the counterweights.

The formula

There is no formula for a north star. What is worth writing down is why it cannot serve as an experiment's decision rule, which is an arithmetic point rather than a philosophical one.

Why it cannot decide a test
n ∝ σ² / Δ²

A company-level metric has a large σ and any single experiment moves it by a tiny Δ. The required sample exceeds the whole user base.

The chain that should exist
primary metric → intermediate outcome → north star

Each link needs an estimated transfer rate. Most programmes assert this chain rather than measuring it — see proxy metric.

A composite, when used
north star = Σ wᵢ · componentᵢ

Weights fixed and public. A composite adjusted quietly after a bad quarter has stopped measuring anything.

The gaming check
watch counter-metrics alongside

Churn, complaints, unsubscribes. The cheap routes to moving a headline number usually show up here first.

Worked example

A streaming service uses hours watched per subscriber per month as its north star, currently 34.2. A team wants to know whether their autoplay-next-episode experiment, which lifted session length by 4%, can be evaluated against it — and what the change is worth at company level.

North star
34.2 hours per subscriber per month
Month-to-month variation
SD 11.6 hours across subscribers
Experiment's effect on session length
+4.0%
Implied effect on monthly hours
≈ +0.4 hours (1.2%)
Sample needed to detect that directly
≈ 1.4 million subscribers per arm
Actual subscriber base
890,000

The north star cannot evaluate this experiment — detecting the effect on it directly would take more subscribers than the company has.

This is the ordinary case rather than an extreme one, and it is why the two kinds of metric must not be conflated. The experiment is judged on session length, which moves detectably in two weeks; the north star is the thing session length is believed to feed, and the belief is what needs periodic validation rather than per-experiment testing. The second half of the question — what is it worth — is where the chain earns its keep: 4% on session length maps to roughly 1.2% on monthly hours, which is a defensible forecast only if that mapping has been estimated from past shipped changes rather than assumed. There is also a gaming caution sitting in plain sight here. Autoplay increases hours watched and can do so by making it harder to stop rather than by delivering more value, and hours watched alone cannot distinguish those. Whether this ships should depend on the churn and satisfaction counter-metrics as much as on the headline.

Common misconceptions

Every experiment should be evaluated against the north star metric.
Almost none can be. A company-level metric is too slow and too noisy for a two-week test — detecting a single experiment's contribution usually requires more users than exist. Experiments are decided on a sensitive primary metric; the north star is what the programme is steering towards in aggregate.
A north star metric means the company only needs one metric.
It needs one metric with authority, plus counter-metrics that constrain it. Any measure given organisational weight gets optimised, including through routes nobody intended, and a headline number without counterweights makes those routes attractive. Engagement paired with churn and complaint rates is far harder to game than engagement alone.
Revenue is the obvious north star.
Revenue is an outcome rather than a driver, and it responds to pricing and discounting far faster than to product quality — which makes it easy to move in ways that damage the business. Most durable north stars measure delivered customer value that revenue follows from, which is why booked nights and weekly active teams are the canonical examples rather than the money.

Frequently asked questions

How does a north star metric differ from a primary metric?
By scope and speed. A primary metric decides a single experiment and must move detectably within its duration, so it is nominated per test and chosen for sensitivity. A north star operates at company scale over quarters and is chosen for stability and alignment. Using a north star to judge an experiment fails on arithmetic — the sample required usually exceeds the entire user base.
How do I connect experiments to the north star?
Build an explicit chain: each experiment's primary metric feeds an intermediate outcome, which feeds the north star, with an estimated transfer rate at each link. Estimate those rates from shipped changes rather than asserting them — hold a long-horizon readout on a sample of what you release and look back at what actually moved. A chain that has never been measured is a story, and it will overstate the value of everything.
How often should a north star metric change?
Rarely — its value comes from being stable enough that teams can orient around it, and one that changes annually provides no alignment at all. Legitimate reasons to revisit include a genuine shift in business model or discovering the metric is being gamed in a way counter-metrics cannot catch. What is not legitimate is changing it because the current number is disappointing.

Related terms

  • Guardrail metric

    A metric that can veto a launch but never justify one, and why that asymmetry is the whole point.

  • Leading and lagging indicators

    The trade between knowing something useful and knowing it in time — and the predictive claim that has to be earned.

  • Primary metric

    The one number the decision hangs on — nominated before the data arrives, which is the entire point.

  • Proxy metric

    A stand-in for the outcome you cannot wait for — and the correlation it rests on is an assumption, not a finding.

Calculate it

  • A/B test sample size

    Size a two-proportion experiment before you launch, then read the lift and its interval once it lands.

Knowing the term is the easy part

Applying it to a live measurement problem is the part that goes wrong. If you are designing an experiment, reading a result you do not trust, or trying to work out what your marketing actually caused, that is the work we do.

References