Bell Statistics

Metrics and measurement

Choosing what an experiment is judged on: the one metric that decides, the ones that protect you, and the ones that only look like progress.

13 terms

An experiment is only as good as the number it is judged on, and that number is chosen long before any data arrives. Most of the ways a well-run test produces a bad decision trace back to this stage — a metric too noisy to move, a metric that improves while the business gets worse, or a dozen metrics watched at once so that something was always going to look significant.

The structure that survives contact with reality is a small one. One primary metric that the decision hangs on, chosen in advance and not negotiable afterwards. A few guardrail metrics that must not degrade, watched for harm rather than for improvement. Everything else is diagnostic — useful for understanding what happened, and not evidence of anything on its own.

The hard cases are the metrics that take too long. Retention, lifetime value and churn are the outcomes anyone actually cares about, and none of them can be read inside a two-week experiment. That is what a proxy metric is for, and the entry says plainly what has to be true for one to be trustworthy — a condition that is checked far less often than it is assumed.

The rest of this group is about how a metric behaves statistically. Whether it is binary, a count or a measured quantity decides which test applies; whether its denominator is random decides whether the standard error is even correct. Each entry links to the calculator that computes it.

Terms in this group

  • Cannibalization

    Moving demand and calling it growth — the failure that only a total-level metric can see.

  • Conversion rate

    Three arbitrary choices wearing a percentage sign — and the reason two teams report different rates for the same week.

  • Guardrail metric

    A metric that can veto a launch but never justify one, and why that asymmetry is the whole point.

  • Halo effect

    The gains that land where nobody was measuring — cannibalization's mirror image, and the reason good work looks flat.

  • Hangover effect

    The cost of relearning, mistaken for a worse product — and the reason a good change can lose its first week.

  • Leading and lagging indicators

    The trade between knowing something useful and knowing it in time — and the predictive claim that has to be earned.

  • Metric types

    Binary, count or continuous — the classification that quietly decides which test is correct and how much traffic you need.

  • North star metric

    A coordination tool for the company, not a decision rule for an experiment — and confusing the two is the usual mistake.

  • Novelty effect

    Curiosity, measured and mistaken for improvement — and the reason a strong week-one result is the least trustworthy kind.

  • Primary metric

    The one number the decision hangs on — nominated before the data arrives, which is the entire point.

  • Proxy metric

    A stand-in for the outcome you cannot wait for — and the correlation it rests on is an assumption, not a finding.

  • Ratio metric

    When the denominator is random too, the ordinary standard error is wrong — and the interval it produces is too narrow.

  • Secondary metric

    Explains the result rather than deciding it — and the moment one gets promoted, the experiment stops meaning what it claims.

A/B Testing at Bell Statistics

We help teams pick metrics that can actually move inside an experiment, and build the guardrails that stop a win in one number being a loss everywhere else. See how we work.