Bell Statistics

Bayesian methods

Reasoning about effects as probability distributions rather than reject-or-not decisions: what that buys, what it costs, and what it does not fix.

7 terms

Frequentist testing asks whether the data would be surprising if nothing were going on. Bayesian methods ask a different question — given what we believed beforehand and what we have now seen, what should we believe about the effect? The answer is a whole distribution rather than a verdict, which supports statements like "an 88% chance the variant is better" that a p-value cannot make.

That reframing is genuinely useful and it is routinely oversold. A posterior distribution makes decision-theoretic reasoning natural: expected loss puts a cost on shipping the wrong arm, which is closer to the actual business question than a significance threshold. Where it is oversold is as a remedy for the peeking problem — Bayesian quantities are not automatically safe under continuous monitoring, and treating them as though they were reintroduces the error it was meant to remove.

Everything turns on the prior. Where real information exists — historical effect sizes, a well-understood market, the domain knowledge that goes into marketing mix modelling — a prior is an asset, and Bayesian methods are the natural framing. Where it does not, an uninformative prior mostly reproduces the frequentist answer in different notation, at the cost of a decision nobody examined.

These entries say what each quantity means, how to read it honestly, and where the framework earns its place against the alternative.

Terms in this group

  • Bayes' theorem

    Evidence updates a belief, it does not replace one — and ignoring the base rate is how a strong test gives a weak conclusion.

  • Bayesian A/B testing

    Friendlier output, the same underlying evidence — and it does not fix peeking, which is why most teams adopt it.

  • Credible interval

    The interval that means what everyone thinks a confidence interval means — conditional on a prior somebody chose.

  • Expected loss

    How much a wrong decision would cost, in the units of the metric — the closest any of these numbers gets to a business answer.

  • Posterior distribution

    The whole distribution of what the effect might be — which is why Bayesian reports can answer questions a p-value cannot.

  • Prior distribution

    What you believed before the data — an asset when it carries real information, and a hidden assumption when it does not.

  • Probability to be best

    The chance an arm is the winner — silent on the margin, and it splits awkwardly across near-identical variants.

A/B Testing at Bell Statistics

We help teams choose the framework that fits the decision rather than the one that produces the friendlier-sounding number. See how we work.