A revenue-per-user metric has a variance of 8,281 (a standard deviation of £91). An analyst decomposes it to see what could be removed, using a pre-experiment window and a day-of-week decomposition.
- Total variance
- 8,281 (SD £91)
- Attributable to stable user differences
- 3,560 (43%)
- Attributable to day-of-week and seasonality
- 660 (8%)
- Attributable to the top 1% of orders
- 2,400 (29%)
- Residual
- 1,661 (20%)
- Sample needed for a 3% MDE, unadjusted
- 184,000 per arm
CUPED on the pre-period covariate removes the 43%; winsorising at the 99th percentile removes most of the 29%. Together the required sample falls from 184,000 to about 60,000 per arm.
Two thirds of this metric's variance had nothing to do with the experiment. The stable user differences are the largest single block and the easiest to remove, because a user's own history predicts it — that is exactly what CUPED subtracts. The top 1% of orders is the second block and the more delicate one: winsorising removes it, and it also changes what you are measuring, since a treatment that genuinely produces very large orders would now be partly invisible. That is an acceptable trade for most conversion-oriented tests and a bad one for a test aimed at high-value customers, which is why the decision belongs in the design document rather than in the analysis. The threefold reduction in sample is what a fortnight instead of six weeks looks like.