Research essays on how marketing measurement actually works — and where it quietly fails. Written from the primary literature: causal inference, geo experiments, Bayesian media mix modeling, and the design of the next experiment.
Part I
Before any method can help, it pays to understand why the naive answers are wrong. Observational data flatters advertising, and untangling cause from correlation takes deliberate machinery.
People who browse more see more ads and do more of everything else online. That single fact can inflate observational estimates of ad effectiveness by orders of magnitude.
A working tour of the identification strategies that stand between you and a causal claim when you cannot randomize: adjustment, instruments, panels, discontinuities, and sensitivity analysis.
One regression fits one causal question, but the output table tempts you to read every coefficient as an effect. Only the exposure's is identified — the rest mix confounding, mediation, and precise nonsense.
Attribution splits credit for conversions that happened; incrementality asks which would not have happened without the ad. The gap is largest exactly where budgets concentrate — retargeting and branded search.
Part II
The workhorse designs for measuring incrementality without a full randomized experiment — and the assumptions each one quietly leans on.
Time-based regression turns a geo split into a causal estimate of ad-driven lift. It works — right up until the pre-period relationship between markets stops holding.
When a marketplace makes user-level A/B tests biased, randomize the whole system on and off over time instead. Carryover sets the switching interval — and long adstock is what breaks it.
When only one market gets the treatment, a weighted blend of untreated markets can stand in for the counterfactual — if you respect the method's bias theory, placebo inference, and requirements.
Two-way fixed effects breaks when treatment rolls out in waves and effects differ across markets or build over time. Group-time average treatment effects rebuild DiD from clean comparisons — and tell you exactly what you're averaging.
Checking that pre-treatment leads are insignificant feels like proof of parallel trends. It isn't: the test is underpowered against the trends that bias you most, and conditioning on passing it distorts the estimate itself.
Fit a state-space model to the pre-period, forecast the world where the campaign never launched, and read the lift as the gap — with honest uncertainty that widens over time.
Part III
Media mix models encode strong assumptions about carryover, saturation, and selection. These essays cover the modeling choices that matter and the mistakes that survive peer review.
Adstock and saturation are where a media mix model's causal claims live. How the carryover and Hill-shape transforms work, why their parameters are hard to identify, and what priors buy you.
An MMM's control coefficients are adjustment terms, not effects — and a downstream "control" like branded search quietly strips an upper-funnel channel's credit. Good and bad controls in a media mix model.
Budgets need long-run effects; experiments observe short-term proxies. A single surrogate rarely satisfies Prentice's criterion — but an index of many short-term signals, validated against occasional truth, can bridge the gap.
An observational MMM coefficient is not causal — spend is endogenous and collinear. A geo lift test identifies what the likelihood cannot, and calibration carries it back in as a prior.
Carryover doesn't just complicate models — it contaminates experiments. How long to wash out, how long to measure, and when a lagged effect makes back-to-back tests read each other's answers.
A field guide to the failure modes that quietly invalidate applied statistical work — from confounded regressions and the R-squared trap to the garden of forking paths.
Select the best of hundreds of model runs and its performance is almost surely exaggerated. Type S/M errors, the winner's curse, and why partial pooling beats correction-by-penalty.
A significant result is both weak evidence and easy to manufacture. Researcher degrees of freedom inflate false positives — and a p just under 0.05 leaves the null with about a one-in-four chance of being true.
Prior predictive checks catch priors that imply impossible outcomes; simulation-based calibration verifies the posterior computation itself — two checks that run before a single real observation.
R-hat, ESS, and divergences each point to a different disease — and only one is cured by running longer. An escalation ladder for failed Bayesian fits, ending with SMC as the multimodality second opinion.
Part IV
Measurement is a loop, not a report. Information theory tells you which experiment to run next; bandits tell you how to earn while you learn.
The expected-information-gain criterion is older than most measurement teams realize. From Lindley's 1956 measure to variational estimators and amortized design policies.
Encode a holdout test as a design vector, treat the MMM posterior as the prior, and the value of the experiment becomes a computable quantity — expected information gain.
Three adaptive frameworks, three different goals: learn the model, find the optimum, or earn while learning. Choosing the wrong one wastes budget in predictable ways.
Sample from the posterior, act as if the sample were true, repeat. Why this one-line algorithm earns near-optimal regret, and what breaks when feedback is delayed or non-stationary.
Closing the loop: a standing program of designed experiments that keeps a response model current — including the channel-interaction effects one-off tests never identify.
The framework puts these ideas to work: causal identification, calibrated Bayesian MMMs, and experiment design in one loop.