Every causal question is a question about something that did not happen. Did the campaign drive those sales — that is, would the sales have occurred anyway? Did the feature retain those users, or were they the kind who stay? The quantity you want is the difference between what happened and what would have happened otherwise, and the second half is never observed. Causal inference is the collection of strategies for constructing a credible stand-in for it.
The gold standard constructs it by force. In an A/B test, randomization creates a group that is identical to the treated group in expectation, so the control's outcome is a direct estimate of the counterfactual — and crucially it balances the variables nobody measured as well as the ones they did. When you can randomise, do; the assumptions are minimal and the argument is short.
When you cannot, the counterfactual has to be modelled, and each method is a different bet about what makes a valid substitute. Difference-in-differences uses an untreated group's change over the same period, betting that both would have moved in parallel. Synthetic control builds a weighted blend of untreated units that tracked the treated one before the intervention. Propensity score matching pairs treated units with observationally similar untreated ones, betting that similarity on measured variables implies similarity on everything relevant. Instrumental variables exploit something that shifts treatment without touching the outcome directly.
What they share is that the bet is an assumption rather than a fact, and it cannot be verified from the data that rely on it. Parallel trends before the intervention is evidence for parallel trends after it, not proof. Balance on observed covariates says nothing about unobserved ones. This is the honest difference between an experiment and an observational estimate: one buys its causal claim with design, the other with an argument. A good causal analysis therefore states its assumption prominently and tests how much the conclusion would change if it were wrong.
In marketing this is not an academic distinction, because most of the important questions cannot be randomised at the user level. Brand advertising reaches households, not cookies. Pricing changes are visible. A TV flight cannot be hidden from half the country. The practical toolkit is therefore geo experiments, which recover randomisation at the regional level, and marketing mix modelling, which models the counterfactual across all channels at once. Which to reach for is the subject of popular tools for accurate campaign impact analysis.
The potential outcomes notation makes the problem exact: two outcomes are defined for every unit and only one is ever observed.
A retailer runs a four-week regional radio campaign in six markets. Sales in those markets rose 8.2% year on year. The CMO wants to know how much of that the radio caused, and three analyses are available.
- Naive: treated markets, year on year
- +8.2%
- Untreated markets over the same period
- +5.1%
- Difference-in-differences estimate
- +3.1%
- Synthetic control estimate
- +2.4% (95% CI 0.9% to 3.9%)
- Pre-period fit of the synthetic control
- RMSE 0.7% over 18 months
The naive figure overstates the effect by roughly three and a half times. The two counterfactual methods agree on something between 2% and 3%.
The 8.2% is not wrong as a description; it is wrong as an answer, because it silently assumes the counterfactual was zero growth — and the untreated markets show that assumption was off by five points. Difference-in-differences fixes that by borrowing the untreated trend, at the cost of assuming the six treated markets would have moved in parallel. Synthetic control is more defensible here because the six markets were chosen for radio availability rather than at random, so they are probably not an average sample: it builds a weighted blend of untreated markets that tracked them closely for eighteen months, and the 0.7% pre-period fit is the evidence for that. The two methods agreeing is worth more than either alone, and the honest headline is 2-3% with the assumption named.