Bell Statistics

Statistics glossary

Forty terms from experimentation, causal inference and marketing measurement, defined properly. Each one gets its own page: what it means, the formula, a worked example, and the misreadings that cost people money.

Most statistics glossaries define a term using four other terms you also do not know. These do not. Every entry opens with a definition you can read out loud in one breath, then explains the same idea again at length, shows the arithmetic on real numbers, and finishes with the specific ways the term gets misused — because in practice the damage is almost never done by people who have never heard of a p-value. It is done by people who have heard of one and think it is the probability the result is a fluke.

The terms are grouped by the job you are doing rather than by statistical family. If you are running an experiment and something looks wrong, start with experimentation — sample ratio mismatch is the first thing to rule out. If you are trying to read a result somebody handed you, start with inference: p-value, confidence interval and statistical power between them explain most of what a results table is claiming.

If the question is whether your marketing actually caused anything, that is causal inference and modelling — incrementality, geo experiments and marketing mix modelling are the three tools most companies end up choosing between, and the entries say plainly what each one can and cannot tell you.

Where a term has a calculator, the entry links to it, so you can go from the definition to a number without leaving the site. We build measurement systems for a living — see our A/B testing work or the case studies — and this is the vocabulary those projects run on.

Running experiments

The vocabulary of A/B testing: what you fix before the test starts, what you watch while it runs, and the diagnostics that tell you the result is not trustworthy.

  • A/B testing

    A randomised experiment on live traffic — and randomisation is the only part that makes it causal.

  • CUPED

    Use pre-period data to strip out noise — the same test, 20-50% less traffic, no change to the estimate.

  • Guardrail metric

    A metric that can veto a launch but never justify one, and why that asymmetry is the whole point.

  • Minimum detectable effect

    The smallest lift you designed to catch — a business decision that gets outsourced to statistics.

  • Randomization

    The one mechanism that makes a comparison causal — and the four ways it silently fails.

  • Sample ratio mismatch

    The split is wrong, so the randomisation is broken — the cheapest and most decisive check you can run.

  • Sample size

    Four inputs, one number, fixed before the test runs — and the square law that makes small effects expensive.

  • Sequential testing

    How to look at a running test without breaking it — and what peeking costs when you do not.

Reading a result

What a significance test actually claims, what it does not, and the four numbers — p, alpha, power, effect size — that decide whether a finding means anything.

  • Confidence interval

    The range your data can actually support, and why it answers the business question a p-value cannot.

  • Effect size

    How big the difference is — the number a p-value throws away and a business case runs on.

  • Multiple comparisons

    Test enough things and something will look significant — the arithmetic, and the four fixes.

  • Null hypothesis

    The assumption every test argues against, and the reason you can never prove two things are the same.

  • P-value

    How surprising your data would be if there were no effect — and the four things it is constantly mistaken for.

  • Significance level

    The false-positive rate you agree to in advance — a budget, and one most teams overspend.

  • Statistical power

    The probability your test finds a real effect — and why most tests that report nothing were never able to.

  • Statistical significance

    A verdict about evidence, not about importance — and the difference is where most bad decisions live.

  • Type I error

    The false positive — and the one whose cost lands months later, on a roadmap built around noise.

  • Type II error

    The false negative — invisible by nature, which is why nobody counts how many good ideas it kills.

Cause and effect

Methods for answering "did this cause that?" when you cannot randomise, and the biases that make an observed difference look like a causal one.

  • Causal inference

    Estimating what an action caused by reconstructing what would have happened without it.

  • Confounding variable

    A common cause of both variables — the reason a strong, stable correlation can mean nothing.

  • Difference-in-differences

    Subtract the untreated group's change from the treated group's — and everything rests on parallel trends.

  • Geo experiment

    Randomise regions instead of users — the way to test marketing that cannot be hidden from a person.

  • Incrementality

    The conversions that would not have happened anyway — and the gap between that and what platforms report.

  • Propensity score matching

    Pair like with like on the probability of being treated — and hope nothing important went unmeasured.

  • Selection bias

    When who ends up in the data is not who you meant to study — and more data makes it worse.

  • Synthetic control

    Build the comparison group instead of finding one — the method for when you have one treated unit.

Models and relationships

Regression, marketing mix modelling and the machinery underneath them — including the failure modes that make a model fit beautifully and predict badly.

  • Adstock

    Advertising does not stop working the week it stops running — and this is how models say so.

  • Correlation

    How tightly two variables move together — bounded, unitless, and silent about cause.

  • Diminishing returns

    The tenth million does less than the first — and why average ROAS is the wrong number to budget on.

  • Marketing mix modelling

    One regression across every channel, built on aggregate data — no tracking, and strong assumptions.

  • Multicollinearity

    When predictors move together the model cannot separate them — good predictions, meaningless coefficients.

  • Overfitting

    A model that memorised the noise — excellent on the data it saw, useless on the data it will meet.

  • R-squared

    Share of variance explained — the most quoted and most over-interpreted number in any model output.

  • Regression analysis

    Fit a line through the data — and the phrase 'holding everything else fixed' is where the trouble starts.

Data and distributions

The building blocks. Spread, error, shape and size — the quantities every method above is ultimately built out of.

  • Central limit theorem

    Why averages go bell-shaped even when the data do not — the result that makes ordinary tests work.

  • Normal distribution

    The bell curve — and why your skewed revenue data usually does not break the test anyway.

  • Outlier

    The extreme value that decides your result — and why the rule for handling it must precede the data.

  • Standard deviation

    How spread out the data are, in the data's own units — and the reason some metrics need vastly more sample.

  • Standard error

    How much your estimate would move if you ran the study again — precision, not spread.

  • Variance

    Spread in squared units — awkward to read, and the quantity every sample-size formula is built on.

Frequently asked questions

Who is this glossary for?
Anyone who has to read, run or argue about an experiment without a statistics degree — product managers, growth and marketing teams, analysts early in their career, and engineers who have inherited an experimentation platform. Each entry assumes no prior notation and introduces any symbol it uses.
How is each entry structured?
A short definition first, so you can leave immediately if that is all you needed. Then a plain-English explanation, the formula where the term is a quantity, a worked example with real numbers, the common misconceptions, and links to the related terms and to any calculator that computes it.
Why are the formulas written without LaTeX?
Because a formula you can select and paste into a spreadsheet or a message to a colleague is more useful than one rendered as an image or a font you do not have. Everything here uses ordinary characters, so it copies as text.
A term I need is not here. Can you add it?
Probably, and we would like to know which one. The list is weighted towards experimentation, causal inference and marketing measurement because that is the work we do, so there are gaps in areas like time-series forecasting and survey methodology. Get in touch and tell us what you were looking for.
Can I link to or quote these definitions?
Yes. Link to the individual term page rather than this index — each one has a stable URL that will not change — and a citation back to it is appreciated but not required. Each entry also lists the textbooks and papers behind it if you need a primary source.

Knowing the words is not the same as knowing the answer.

Most measurement problems are not vocabulary problems. They are a metric that moves for the wrong reason, a test that never had the power to detect what you were looking for, or three channels each claiming the same conversion. If any of that sounds familiar, talk to us.