A team plans a test needing 200,000 users per arm and wants the option to stop early. They compare four-look O'Brien-Fleming and Pocock boundaries against a fixed-horizon design, simulating both a true 6% lift and no effect at all.
- Fixed horizon
- 200,000/arm, one look at p < 0.05
- O'BF boundaries
- 0.0001 / 0.004 / 0.019 / 0.043
- Pocock boundaries
- 0.018 at each of four looks
- Real effect: O'BF stops early
- 31% of runs, mean 148,000/arm
- Real effect: Pocock stops early
- 58% of runs, mean 121,000/arm
- Power at full horizon: fixed / O'BF / Pocock
- 80.0% / 79.2% / 74.6%
Pocock stops early far more often and gives up 5.4 points of power at the horizon. O'Brien-Fleming stops early half as often and costs almost nothing.
The last row is the decision. Pocock's higher early-stopping rate looks attractive until you notice what it costs on the experiments that do not stop — 74.6% power against 80%, which means one experiment in twenty that would have found a real effect now misses it. Since most experiments run to the horizon, that penalty applies to the majority while the benefit applies to the minority. O'Brien-Fleming inverts that: it stops early only when the effect is large enough to be obvious, and preserves essentially full power for everything else. For product experimentation, where continuing a test costs calendar time rather than anything serious, that is the right trade. Pocock makes sense when continuing is genuinely expensive or harmful — a clinical trial where patients are receiving an inferior treatment — which is a situation product teams rarely face.