Reading a result
What a significance test actually claims, what it does not, and the four numbers — p, alpha, power, effect size — that decide whether a finding means anything.
A results table is a set of claims in a compressed notation, and almost every expensive mistake in experimentation comes from reading one of those claims as something adjacent to what it says. A p-value is not the probability the result is wrong. A non-significant result is not evidence of no effect. A significant one is not evidence the effect is big enough to ship. Each of those is a different error and each has its own entry here.
The four numbers are connected, which is the part that is rarely taught together. Statistical power, the significance level, the effect size you care about and the sample size you can afford form a system with three degrees of freedom: fix any three and the fourth is determined. Test planning is choosing which one to sacrifice, and doing it deliberately rather than discovering it afterwards.
The single most useful habit this group argues for is reporting a confidence interval next to every p-value. The p answers a yes-or-no question about a hypothesis nobody believed; the interval answers the question the business is asking, which is how much. They are computed from the same numbers and never disagree — but only one of them tells you whether the result is worth acting on.
