
Five Common Blind Spots in Marketing Data
In today’s privacy-focused era, the different attribution models create many blind spots for marketing analysts and decision makers. However MMM & Geo Tests can help.

Heterogeneous treatment effects exist when a change helps some units more than others, or helps some and harms others. The conditional average treatment effect is the effect within a defined subgroup, and estimating it well is much harder than estimating the overall average.
CATEThe average treatment effect answers what happens if you treat everyone. It rarely describes what happens to anyone. A simplified onboarding flow can help users arriving cold and hinder those who already know the product; a discount can move price-sensitive buyers and give margin away on the rest. When effects differ across units like that, they are heterogeneous, and the effect within a particular subgroup is the conditional average treatment effect — the CATE.
Why this matters practically is that heterogeneity creates options the average conceals. If a change helps one segment by four points and harms another by two, shipping it to the first alone beats shipping it to everyone, and the overall figure would never have suggested that. A great deal of the value in a mature experimentation programme comes from finding those splits rather than from finding universally good changes, which get rarer as a product matures.
The difficulty is that looking for heterogeneity is one of the easiest ways to find things that are not there. Slice a null result by device, country, tenure, channel and plan, and something reverses by chance — the multiple comparisons problem with a large and undocumented number of paths. Interaction effects also need roughly four times the sample of main effects to detect reliably, so most experiments are badly underpowered for the question and will produce noise that looks like insight.
Two disciplines make the search credible. Pre-specify the segments you will examine, before the data exists, and report the interaction test rather than the segment effects on their own — a difference between segments is the claim, and it needs its own p-value. Where exploratory search is genuinely wanted, causal machine learning methods such as causal forests are built for it and handle the multiplicity honestly, at the cost of requiring substantially more data and more care than a pivot table.
Whatever the method, a heterogeneity finding is a hypothesis until it is confirmed on new data. The segment that looked different is the segment that looked most different out of however many were examined, so its estimated effect carries the same winner's curse inflation as any selected result. Re-testing on the nominated segment as the primary population is what converts it into something worth building a targeting rule on.
The conditional estimand, the interaction test that establishes it, and the sample penalty that makes it hard.
CATE(x) = E[ Y(1) − Y(0) | X = x ]The effect within the subgroup defined by covariates X. The ATE is its average over the population.
β₃ in Y = β₀ + β₁·treated + β₂·segment + β₃·(treated × segment)β₃ is the difference in effect between segments. This is the quantity to test, not the two effects separately.
n_interaction ≈ 4 × n_main effectA difference of differences, so noise compounds — see the sample size calculator.
1 − ( 1 − α )^k over k segmentationsFive segmentations at 5% gives a 23% chance of a spurious interaction.
A pricing page test returns an overall effect indistinguishable from zero. Two analyses follow: one pre-registered split by company size, and an exploratory sweep across twelve other segmentations.
The pre-registered split is a real finding. The two significant results from twelve exploratory slices are close to what chance predicts.
The contrast between the two halves is the whole lesson. The company-size split was nominated before the data existed, the experiment was powered with it in mind, and its interaction test at p = 0.004 is a genuine result — a change that helps small customers and harms large ones, which is directly actionable through targeting. The exploratory sweep produced two significant interactions from twelve attempts, against 0.6 expected by chance, which is unremarkable and would not survive a correction. Reporting those two alongside the first as though they were equivalent findings is how a programme accumulates targeting rules that do nothing. Note also that the overall ATE of +0.3% is not wrong — it is the correct answer to a question nobody should have asked here, since a change with effects of +6.2% and −4.1% in different directions was never sensibly evaluated as a single number.

In today’s privacy-focused era, the different attribution models create many blind spots for marketing analysts and decision makers. However MMM & Geo Tests can help.


Understanding why the uplift seen in A/B tests often differs from real-world outcomes is essential. Explore key reasons behind this discrepancy and insights on how to manage expectations effectively.

Applying it to a live measurement problem is the part that goes wrong. If you are designing an experiment, reading a result you do not trust, or trying to work out what your marketing actually caused, that is the work we do.