Research essays on how marketing measurement actually works. Most of its failures are quiet ones, so they get as much room here as the methods do. Everything is written from the primary literature: causal inference, geo experiments, Bayesian media mix modeling, and the design of the next experiment.
Part I
Before any method can help, it pays to understand why the naive answers are wrong. Observational data flatters advertising, and untangling cause from correlation takes deliberate machinery.
People who browse more see more ads and do more of everything else online. That single fact can inflate observational estimates of ad effectiveness by orders of magnitude.
A working tour of the identification strategies that stand between you and a causal claim when you cannot randomize: adjustment, instruments, panels, discontinuities, and sensitivity analysis.
One regression fits one causal question, but the output table tempts you to read every coefficient as an effect. Only the exposure's is identified. The rest mix confounding, mediation, and precise nonsense.
Attribution splits credit for conversions that happened. Incrementality asks which of them would not have happened without the ad. The gap is largest exactly where budgets concentrate, in retargeting and branded search.
Randomizing media identifies the total effect and never the direct/indirect split. And "proportion mediated" is not even a proportion.
Partial adjustment is not attenuation. Under a condition media planning routinely satisfies, controlling for a demand proxy increases the bias. A better proxy increases it more.
You cannot test the no-confounding assumption, so price it instead. Frequentist bounds ask how strong a confounder would have to be; the Bayesian route makes it a parameter. The E-value does not apply to an ROI.
Row-wise LOO on a weekly panel trains on the weeks after the one it held out, so it scores an interpolation nobody performs. The fold is wrong, not the estimator, and the diagnostics thrown away with the scalar are worth more than it.
Posterior curvature has two owners. Splitting it gives a number that is not the parameter count: how many parameters the likelihood actually determines, and which ones are the prior wearing a posterior.
A model cannot fail a check on a quantity its own parameters reproduce. The check becomes evidence only against a statistic the fit was never asked to match.
The Bayesian workflow fits many models; pre-specification commits before looking. They reconcile under one condition, and the spread across the specs you happened to fit is not a bracket on the truth.
One fit already contains every prior in a neighbourhood of the one you used. Power-scaling recovers them from draws you have, which retires the paragraph that says three priors were tried and results were stable.
Rank calibration certifies the sampler and says nothing about whether the model was worth fitting. The same simulation loop answers the second question, and marginal ranks cannot see the failure an MMM has most often.
The front-door escape hatch is now falsifiable, and it fails on ad data. The randomization that works was already performed by the platform's pacing system. Whether it was logged for you is the open question.
Part II
The workhorse designs for measuring incrementality without a full randomized experiment. Each one leans quietly on an assumption, and these essays name it.
Time-based regression turns a geo split into a causal estimate of ad-driven lift. It works right up until the pre-period relationship between markets stops holding.
When a marketplace makes user-level A/B tests biased, randomize the whole system on and off over time instead. Carryover sets the switching interval. Long adstock is what breaks it.
Randomizing users is not enough to compare ad content, because the algorithm re-sorts the audience per creative. Lift tests survive this. Creative A/Bs do not, and a biased-but-precise readout will overrule your geo holdout.
When only one market gets the treatment, a weighted blend of untreated markets can stand in for the counterfactual. That holds only if you respect the method's bias theory, its placebo inference, and its data requirements.
Two-way fixed effects breaks when treatment rolls out in waves and effects differ across markets or build over time. Group-time average treatment effects rebuild DiD from clean comparisons. They also tell you exactly what you're averaging.
Checking that pre-treatment leads are insignificant feels like proof of parallel trends. The test is underpowered against exactly the trends that bias you most, and conditioning on passing it distorts the estimate itself.
Fit a state-space model to the pre-period, forecast the world where the campaign never launched, and read the lift as the gap between them. The uncertainty widens honestly as the forecast runs on.
Part III
Media mix models encode strong assumptions about carryover, saturation, and selection. These essays cover the modeling choices that matter and the mistakes that survive peer review.
Adstock and saturation are where a media mix model's causal claims live. How the carryover and Hill-shape transforms work, why their parameters are hard to identify, and what priors buy you.
Diminishing returns and declining effectiveness fit the same series and imply opposite budgets. The adstock transform is what makes them indistinguishable.
The average fit stays clean, and even improves, as you sum finer spend into coarser totals. Meanwhile the marginal return the budget optimizer reads drifts quietly away from the truth.
Holdout MAPE scores the sum. ROI depends on the split, and the split is where all the identification lives. What to put in the contract instead.
An experiment that calibrated your model's prior cannot also be cited as independent evidence that the model is right. The framework's own triangulation report currently calls this "convergent evidence from two independent methods."
Trend flexibility is a causal hyperparameter, and it gets tuned as if it were a fit hyperparameter. A tighter baseline gives you a narrower media interval centered in the wrong place, and no diagnostic in the output warns you.
Your model says spend buys a decaying asset. The optimizer allocates a window total and hands the timing question to a taste menu. Two defects you can quantify from your own fit.
Ranking a continuous budget simplex by a finite posterior sample biases the winner's own reported value upward. The framework already prices decision instability, and this is the next thing to price.
Staying inside every channel's own historical spend range does not mean the fitted curve still applies once you act on it. The framework's own honest extrapolation check is blind to the failure by construction.
An MMM's control coefficients are adjustment terms rather than effects. A downstream "control" like branded search quietly strips an upper-funnel channel's credit. Good and bad controls in a media mix model.
A sign-guarded price elasticity forecloses one failure mode and leaves the magnitude exactly where an uninstrumented regression puts it. One parameter, two questions, and only one of them was ever answered.
Budgets need long-run effects. Experiments observe short-term proxies. A single surrogate rarely satisfies Prentice's criterion, though an index of many short-term signals, validated against occasional truth, can bridge the gap.
An observational MMM coefficient is not causal, because spend is endogenous and collinear. A geo lift test identifies what the likelihood cannot, and calibration carries it back in as a prior.
Carryover contaminates experiments as thoroughly as it complicates models. How long to wash out, how long to measure, and when a lagged effect makes back-to-back tests read each other's answers.
A field guide to the failure modes that quietly invalidate applied statistical work, from confounded regressions and the R-squared trap to the garden of forking paths.
Select the best of hundreds of model runs and its performance is almost surely exaggerated. Type S/M errors, the winner's curse, and why partial pooling beats correction-by-penalty.
A significant result is both weak evidence and easy to manufacture. Researcher degrees of freedom inflate false positives. A p just under 0.05 leaves the null with about a one-in-four chance of being true.
Ranking variables by p-value across regressions compares sample sizes, noise levels, and spend variation rather than effects. Between two near-null variables the contest is exactly a coin toss. A real comparison needs a statistic for the difference.
A non-significant test cannot show two groups are the same. Silence is not absence. Equivalence margins and TOST on the frequentist side, ROPE probabilities and Bayes factors on the Bayesian side, and why proving nothing costs more than finding something.
Prior predictive checks catch priors that imply impossible outcomes. Simulation-based calibration verifies the posterior computation itself. Both run before a single real observation arrives.
R-hat, ESS, and divergences each point to a different disease, and only one of the three is cured by running longer. An escalation ladder for failed Bayesian fits, ending with SMC as the multimodality second opinion.
Part IV
Measurement runs as a loop. Information theory tells you which experiment to run next, and bandits tell you how to earn while you learn.
The expected-information-gain criterion is older than most measurement teams realize. From Lindley's 1956 measure to variational estimators and amortized design policies.
Encode a holdout test as a design vector, treat the MMM posterior as the prior, and the value of the experiment becomes a computable quantity called expected information gain.
Three adaptive frameworks, three different goals: learn the model, find the optimum, or earn while learning. Choosing the wrong one wastes budget in predictable ways.
Sample from the posterior, act as if the sample were true, repeat. Why this one-line algorithm earns near-optimal regret, and what breaks when feedback is delayed or non-stationary.
A standing program of designed experiments keeps a response model current, including the channel-interaction effects that one-off tests never identify.
Coverage tells you an interval was right often enough. It cannot tell you the forecaster was honest about it. The sixty-year-old fix is a proper scoring rule, and this framework's scorecard doesn't carry one yet.
A second estimation path made a hidden dependency visible: the media priors were quietly offsetting confounding bias. A penalty chosen by cross-validation cannot do that job, because under confounding the coefficient that predicts best is the confounded one.
The framework puts these ideas to work: causal identification, calibrated Bayesian MMMs, and experiment design in one loop.