A significance test tells you whether an effect is distinguishable from zero. Effect size tells you how big it is. The two come apart badly at scale: with enough users, a difference too small to fund is detected with near-certainty, and with too few, a transformative difference goes unproven. Only one of these numbers appears in a business case, and it is not the p-value.
There are two families, and mixing them up is the most expensive vocabulary error in experimentation. Unstandardised effect sizes are in the units of the thing itself: 0.32 percentage points of conversion, £1.40 of revenue per user, eleven seconds of session length. Standardised ones divide by a measure of spread — Cohen's d is a difference in means divided by the pooled standard deviation — which makes effects comparable across metrics measured on different scales, at the cost of being unreadable to anyone outside the analysis.
In product and marketing work the unstandardised version is almost always the right one to report, because the decision is denominated in the same units. "Revenue per user rose by £1.40" supports a decision; "d = 0.06" does not, and quietly invites the reader to look up a benchmark table. Those tables — Cohen's 0.2 small, 0.5 medium, 0.8 large — come from psychology experiments on tens of participants and are wildly out of place online, where a d of 0.02 on a hundred million sessions can be worth millions and would be classed as negligible.
The relative-versus-absolute distinction causes the other half of the confusion, and it is worth being pedantic about. A move from 4.0% to 4.2% is 0.2 percentage points absolute and a 5% relative lift. Both are correct; they sound very different; and a report that says "conversion up 5%" without saying which will be read as the larger one. Absolute is what compounds across surfaces and what the arithmetic produces. Relative is what business cases are written in. Say which.
Effect size is also an input, not just an output. Every sample-size calculation starts from the smallest effect worth detecting — the minimum detectable effect — and that number has to be chosen on business grounds before the test runs. Choosing it afterwards, from the observed result, is how tests end up powered for whatever happened to occur. And beware the observed effect from an underpowered test: because only inflated estimates clear significance, the winners you see are systematically overstated, sometimes by a factor of two.
The unstandardised forms are subtraction and division. The standardised ones divide by spread, which is what makes them comparable across metrics and unreadable to stakeholders.
A checkout experiment on 240,000 sessions per arm moves conversion from 3.10% to 3.19%. Session-level revenue rises from £2.05 to £2.12. The team wants to know whether to ship, and the finance model needs an annual number against 90 million sessions.
- Absolute difference
- +0.09 percentage points
- Relative lift
- +2.9%
- Cohen's h
- 0.0051
- Revenue per session
- +£0.07
- 95% CI on revenue
- +£0.01 to +£0.13
Statistically significant (p = 0.02), negligible by Cohen's benchmarks, and worth roughly £6.3m a year at the point estimate — with the interval spanning £0.9m to £11.7m.
Three effect sizes, three different stories, and only one of them is the decision. Cohen's h of 0.005 would be classed as far below "small" by a benchmark table built for psychology experiments, and it is the least informative number here. The 2.9% relative lift is the one that will get repeated in the meeting. The absolute £0.07 per session is the one the finance model needs, and its interval is the honest version: this is worth somewhere between one and twelve million a year, which is a ship decision and also a reason not to write £6.3m in a plan as though it were a measurement.
- דA small p-value means a large effect.”
- It means a well-measured one. The p-value combines effect size and sample size, so an enormous sample makes trivial effects highly significant and a small sample leaves large effects unproven. The magnitude lives in the estimate and its confidence interval, and neither is recoverable from p alone.
- דCohen's d of 0.2 is small, so an effect that size is not worth having.”
- Those benchmarks come from mid-century psychology studies and travel badly. Online, a standardised effect an order of magnitude below Cohen's "small" can be worth millions a year, because the population is enormous and the change is permanent. Judge magnitude against your own economics, not a table.
- דOur test found a 12% lift, so we should plan for a 12% lift.”
- The point estimate is the middle of a range, and in an underpowered test it is biased upward — only effects that noise inflated reach significance, so the winners you see overstate the truth. Plan against the lower end of the interval, and treat the gap between that and the point estimate as the reason the number in the deck rarely appears in the P&L.