
Five Common Blind Spots in Marketing Data
In today’s privacy-focused era, the different attribution models create many blind spots for marketing analysts and decision makers. However MMM & Geo Tests can help.

A confounding variable is something that influences both the supposed cause and the outcome, creating an association between them that is not causal. It is the reason an observed relationship can be strong, stable, statistically significant and completely misleading.
Two things move together. A third thing was driving both. That third thing is a confounder, and it is the central obstacle to learning anything causal from data you merely observed. The textbook example is ice cream sales and drowning deaths, which rise and fall together because summer causes both. The marketing example is spend and revenue, which rise and fall together because demand causes both — you spend more when people are buying more, and the correlation between spend and sales is therefore guaranteed whether or not the spend works.
What makes confounding dangerous rather than merely inconvenient is that it produces good-looking evidence. The relationship is real in the data, it replicates, it is highly significant, and more data makes the estimate more precise without making it less wrong. There is no diagnostic in a regression output that flags it. Confounding is a property of how the world generated the data, and no amount of scrutiny applied to the data alone will reveal it.
There are two standard responses. Control for it — include it in a regression, match on it, or stratify — which works exactly to the extent that you have measured it well and modelled its relationship correctly. Or design it away, which is what randomization does: assigning treatment by chance breaks the link between the treatment and everything else at once, including the variables you never measured. That is the asymmetry that makes an experiment worth so much more than a regression with controls.
Two related traps are worth naming because both look like careful analysis. A collider is a variable caused by both the treatment and the outcome, and controlling for one *creates* bias rather than removing it — conditioning on "became a customer" when studying whether ad exposure drives spend is a common version. And a mediator sits on the causal path between treatment and outcome; controlling for it removes the very effect you are trying to measure. The rule that follows is that you cannot decide what to control for by looking at the data. It requires a claim about which variable causes which, made before the model is fitted.
In practice, confounding is why platform-reported performance overstates incrementality so consistently. Retargeting is shown to people already close to buying — intent confounds exposure and conversion — so the observed relationship is enormous and the causal one is small. And it is why marketing mix modelling is a discipline rather than a regression: seasonality, promotions, distribution, competitor activity and price all move spend and sales together, and every one of them has to be in the model or accounted for by design.
Confounding has a precise algebraic signature: the omitted variable shows up inside the coefficient you did estimate.
Y = β₀ + β₁·D + β₂·C + εD is the treatment, C the confounder. Fitting Y on D alone does not simply lose β₂ — it corrupts β₁.
E[β̂₁] = β₁ + β₂ · ( Cov(D, C) / Var(D) )The bias is the confounder's own effect times how strongly it correlates with the treatment. Both must be non-zero for confounding to occur.
same sign as β₂ · Cov(D, C)When both are positive — demand raises spend and raises sales — the bias is upward, which is why observational marketing estimates skew optimistic.
Cov(D, C) = 0 for every C, measured or notRandom assignment zeroes the second factor for all confounders at once. This one line is the entire argument for experiments.
An analyst finds that customers who use the mobile app spend 40% more per year than customers who do not, and proposes a campaign to drive app installs on the strength of it. The finding is highly significant across two million customers.
The 40% gap shrinks to 14% once prior spend and tenure are matched, and to about 4% — not distinguishable from zero — in a randomised test.
The confounder here is customer engagement, and it is not subtle once named: people who already buy a lot are more likely to install the app, so the app is a marker of loyalty rather than a cause of it. Matching on prior spend and tenure removes most of the gap, which tells you the confounding was severe. It does not tell you the remaining £31 is causal, because matching only handles what was measured, and the randomised test suggests most of that residue was confounding too. The sequence matters more than the numbers: each step removed bias and each step lowered the estimate, which is the fingerprint of a confounded relationship rather than a real one being refined.

In today’s privacy-focused era, the different attribution models create many blind spots for marketing analysts and decision makers. However MMM & Geo Tests can help.


Understanding why the uplift seen in A/B tests often differs from real-world outcomes is essential. Explore key reasons behind this discrepancy and insights on how to manage expectations effectively.

Applying it to a live measurement problem is the part that goes wrong. If you are designing an experiment, reading a result you do not trust, or trying to work out what your marketing actually caused, that is the work we do.