Metrics and measurement
Choosing what an experiment is judged on: the one metric that decides, the ones that protect you, and the ones that only look like progress.
An experiment is only as good as the number it is judged on, and that number is chosen long before any data arrives. Most of the ways a well-run test produces a bad decision trace back to this stage — a metric too noisy to move, a metric that improves while the business gets worse, or a dozen metrics watched at once so that something was always going to look significant.
The structure that survives contact with reality is a small one. One primary metric that the decision hangs on, chosen in advance and not negotiable afterwards. A few guardrail metrics that must not degrade, watched for harm rather than for improvement. Everything else is diagnostic — useful for understanding what happened, and not evidence of anything on its own.
The hard cases are the metrics that take too long. Retention, lifetime value and churn are the outcomes anyone actually cares about, and none of them can be read inside a two-week experiment. That is what a proxy metric is for, and the entry says plainly what has to be true for one to be trustworthy — a condition that is checked far less often than it is assumed.
The rest of this group is about how a metric behaves statistically. Whether it is binary, a count or a measured quantity decides which test applies; whether its denominator is random decides whether the standard error is even correct. Each entry links to the calculator that computes it.
A/B Testing at Bell Statistics
We help teams pick metrics that can actually move inside an experiment, and build the guardrails that stop a win in one number being a loss everywhere else. See how we work.
