Modern Measurement Research

Research essays on how marketing measurement actually works. Most of its failures are quiet ones, so they get as much room here as the methods do. Everything is written from the primary literature: causal inference, geo experiments, Bayesian media mix modeling, and the design of the next experiment.

Part I

Why Measurement Is Hard

Before any method can help, it pays to understand why the naive answers are wrong. Observational data flatters advertising, and untangling cause from correlation takes deliberate machinery.

≈ 10 min read

Activity Bias: Why Observational Ad Measurement Flatters Itself

People who browse more see more ads and do more of everything else online. That single fact can inflate observational estimates of ad effectiveness by orders of magnitude.

≈ 16 min read

Uncovering Causal Estimates from Non-Experimental Data

A working tour of the identification strategies that stand between you and a causal claim when you cannot randomize: adjustment, instruments, panels, discontinuities, and sensitivity analysis.

≈ 17 min read

The Table 2 Fallacy: When a Coefficient Isn't an Effect

One regression fits one causal question, but the output table tempts you to read every coefficient as an effect. Only the exposure's is identified. The rest mix confounding, mediation, and precise nonsense.

≈ 18 min read

Attribution Is Not Incrementality

Attribution splits credit for conversions that happened. Incrementality asks which of them would not have happened without the ad. The gap is largest exactly where budgets concentrate, in retargeting and branded search.

≈ 24 min read

Randomization Buys You the Total Effect. The Funnel Is Not Included.

Randomizing media identifies the total effect and never the direct/indirect split. And "proportion mediated" is not even a proportion.

≈ 26 min read

Your Demand Proxy Is Not a Control Variable

Partial adjustment is not attenuation. Under a condition media planning routinely satisfies, controlling for a demand proxy increases the bias. A better proxy increases it more.

≈ 15 min read

Two Ways to Price a Confounder You Cannot See

You cannot test the no-confounding assumption, so price it instead. Frequentist bounds ask how strong a confounder would have to be; the Bayesian route makes it a parameter. The E-value does not apply to an ROI.

≈ 16 min read

Your MMM Is Reporting the Wrong Cross-Validation

Row-wise LOO on a weekly panel trains on the weeks after the one it held out, so it scores an interpolation nobody performs. The fold is wrong, not the estimator, and the diagnostics thrown away with the scalar are worth more than it.

≈ 17 min read

How Many Parameters Did Your Data Pay For?

Posterior curvature has two owners. Splitting it gives a number that is not the parameter count: how many parameters the likelihood actually determines, and which ones are the prior wearing a posterior.

≈ 17 min read

A Posterior Predictive Check Cannot See What You Adjusted For

A model cannot fail a check on a quantity its own parameters reproduce. The check becomes evidence only against a statistic the fit was never asked to match.

≈ 19 min read

Fitting Many Models Is Not P-Hacking

The Bayesian workflow fits many models; pre-specification commits before looking. They reconcile under one condition, and the spread across the specs you happened to fit is not a bracket on the truth.

≈ 15 min read

Prior Sensitivity Without Refitting Anything

One fit already contains every prior in a neighbourhood of the one you used. Power-scaling recovers them from draws you have, which retires the paragraph that says three priors were tried and results were stable.

≈ 17 min read

The Simulation Loop Does More Than SBC

Rank calibration certifies the sampler and says nothing about whether the model was worth fitting. The same simulation loop answers the second question, and marginal ranks cannot see the failure an MMM has most often.

≈ 25 min read

The Randomization Is Already in the Auction

The front-door escape hatch is now falsifiable, and it fails on ad data. The randomization that works was already performed by the platform's pacing system. Whether it was logged for you is the open question.

Part II

Quasi-Experimental Methods

The workhorse designs for measuring incrementality without a full randomized experiment. Each one leans quietly on an assumption, and these essays name it.

≈ 15 min read

Geo Experiments: TBR, Power, and the Stationarity You're Assuming

Time-based regression turns a geo split into a causal estimate of ad-driven lift. It works right up until the pre-period relationship between markets stops holding.

≈ 14 min read

Switchback Experiments: Randomizing Time, Not Users

When a marketplace makes user-level A/B tests biased, randomize the whole system on and off over time instead. Carryover sets the switching interval. Long adstock is what breaks it.

≈ 23 min read

The Platform Randomizes After You Do

Randomizing users is not enough to compare ad content, because the algorithm re-sorts the audience per creative. Lift tests survive this. Creative A/Bs do not, and a biased-but-precise readout will overrule your geo holdout.

≈ 16 min read

Synthetic Control, Done Right

When only one market gets the treatment, a weighted blend of untreated markets can stand in for the counterfactual. That holds only if you respect the method's bias theory, its placebo inference, and its data requirements.

≈ 14 min read

Modern Staggered Difference-in-Differences

Two-way fixed effects breaks when treatment rolls out in waves and effects differ across markets or build over time. Group-time average treatment effects rebuild DiD from clean comparisons. They also tell you exactly what you're averaging.

≈ 16 min read

Pretest With Caution: Testing for Parallel Trends

Checking that pre-treatment leads are insignificant feels like proof of parallel trends. The test is underpowered against exactly the trends that bias you most, and conditioning on passing it distorts the estimate itself.

≈ 14 min read

CausalImpact: Counterfactuals from Bayesian Structural Time Series

Fit a state-space model to the pre-period, forecast the world where the campaign never launched, and read the lift as the gap between them. The uncertainty widens honestly as the forecast runs on.

Part III

Models and Their Failure Modes

Media mix models encode strong assumptions about carryover, saturation, and selection. These essays cover the modeling choices that matter and the mistakes that survive peer review.

≈ 16 min read

Carryover and Shape Effects in Bayesian MMM

Adstock and saturation are where a media mix model's causal claims live. How the carryover and Hill-shape transforms work, why their parameters are hard to identify, and what priors buy you.

≈ 24 min read

Saturation or Fatigue? Your MMM Cannot Tell the Difference

Diminishing returns and declining effectiveness fit the same series and imply opposite budgets. The adstock transform is what makes them indistinguishable.

≈ 21 min read

Aggregation Bends the Curve

The average fit stays clean, and even improves, as you sum finer spend into coarser totals. Meanwhile the marginal return the budget optimizer reads drifts quietly away from the truth.

≈ 27 min read

Stop Validating Your MMM With Holdout Error

Holdout MAPE scores the sum. ROI depends on the split, and the split is where all the identification lives. What to put in the contract instead.

≈ 17 min read

Can't Calibrate and Validate on the Same Experiment

An experiment that calibrated your model's prior cannot also be cited as independent evidence that the model is right. The framework's own triangulation report currently calls this "convergent evidence from two independent methods."

≈ 27 min read

The Baseline Ate Your Media Effect

Trend flexibility is a causal hyperparameter, and it gets tuned as if it were a fit hyperparameter. A tighter baseline gives you a narrower media interval centered in the wrong place, and no diagnostic in the output warns you.

≈ 25 min read

Advertising Is a Stock. Your Optimizer Solves a One-Period Problem.

Your model says spend buys a decaying asset. The optimizer allocates a window total and hands the timing question to a taste menu. Two defects you can quantify from your own fit.

≈ 20 min read

The Optimizer's Curse

Ranking a continuous budget simplex by a finite posterior sample biases the winner's own reported value upward. The framework already prices decision instability, and this is the next thing to price.

≈ 21 min read

Your Curve Assumed You Wouldn't Act On It

Staying inside every channel's own historical spend range does not mean the fitted curve still applies once you act on it. The framework's own honest extrapolation check is blind to the failure by construction.

≈ 15 min read

The Table 2 Fallacy in Media Mix Models

An MMM's control coefficients are adjustment terms rather than effects. A downstream "control" like branded search quietly strips an upper-funnel channel's credit. Good and bad controls in a media mix model.

≈ 20 min read

You Modelled One P

A sign-guarded price elasticity forecloses one failure mode and leaves the magnitude exactly where an uninstrumented regression puts it. One parameter, two questions, and only one of them was ever answered.

≈ 17 min read

Measuring the Long Run: Surrogate Outcomes in Media

Budgets need long-run effects. Experiments observe short-term proxies. A single surrogate rarely satisfies Prentice's criterion, though an index of many short-term signals, validated against occasional truth, can bridge the gap.

≈ 13 min read

The Experiment Is the Prior: Calibrating a Media Mix Model

An observational MMM coefficient is not causal, because spend is endogenous and collinear. A geo lift test identifies what the likelihood cannot, and calibration carries it back in as a prior.

≈ 15 min read

Adstock Dynamics and the Timing of Sequential Experiments

Carryover contaminates experiments as thoroughly as it complicates models. How long to wash out, how long to measure, and when a lagged effect makes back-to-back tests read each other's answers.

≈ 13 min read

Common Pitfalls in Statistical Modeling

A field guide to the failure modes that quietly invalidate applied statistical work, from confounded regressions and the R-squared trap to the garden of forking paths.

≈ 15 min read

Hundreds of Models, One Winner: Multiple Comparisons and the Winner's Curse

Select the best of hundreds of model runs and its performance is almost surely exaggerated. Type S/M errors, the winner's curse, and why partial pooling beats correction-by-penalty.

≈ 15 min read

P-Hacking and the Thin Evidence of a P-Value

A significant result is both weak evidence and easy to manufacture. Researcher degrees of freedom inflate false positives. A p just under 0.05 leaves the null with about a one-in-four chance of being true.

≈ 20 min read

A Smaller P-Value Is Not More Signal

Ranking variables by p-value across regressions compares sample sizes, noise levels, and spend variation rather than effects. Between two near-null variables the contest is exactly a coin toss. A real comparison needs a statistic for the difference.

≈ 22 min read

Proving Nothing Happened: Testing for No Difference

A non-significant test cannot show two groups are the same. Silence is not absence. Equivalence margins and TOST on the frequentist side, ROPE probabilities and Bayes factors on the Bayesian side, and why proving nothing costs more than finding something.

≈ 16 min read

Does Your Model Work Before It Sees Data?

Prior predictive checks catch priors that imply impossible outcomes. Simulation-based calibration verifies the posterior computation itself. Both run before a single real observation arrives.

≈ 16 min read

When Sampling Fails

R-hat, ESS, and divergences each point to a different disease, and only one of the three is cured by running longer. An escalation ladder for failed Bayesian fits, ending with SMC as the multimodality second opinion.

Part IV

Designing the Next Experiment

Measurement runs as a loop. Information theory tells you which experiment to run next, and bandits tell you how to earn while you learn.

≈ 16 min read

From Lindley to Deep Adaptive Design: 65 Years of Bayesian Experimental Design

The expected-information-gain criterion is older than most measurement teams realize. From Lindley's 1956 measure to variational estimators and amortized design policies.

≈ 10 min read

A Geo-Holdout Is a Bayesian Experimental Design: Computing Its EIG

Encode a holdout test as a design vector, treat the MMM posterior as the prior, and the value of the experiment becomes a computable quantity called expected information gain.

≈ 14 min read

BED vs. Bayesian Optimization vs. Bandits for Media Experimentation

Three adaptive frameworks, three different goals: learn the model, find the optimum, or earn while learning. Choosing the wrong one wastes budget in predictable ways.

≈ 14 min read

Thompson Sampling: A Practical Tour

Sample from the posterior, act as if the sample were true, repeat. Why this one-line algorithm earns near-optimal regret, and what breaks when feedback is delayed or non-stationary.

≈ 14 min read

Continuous Learning in Media Measurement (with Interaction Effects)

A standing program of designed experiments keeps a response model current, including the channel-interaction effects that one-off tests never identify.

≈ 18 min read

Vendor-Graded Homework

Coverage tells you an interval was right often enough. It cannot tell you the forecaster was honest about it. The sixty-year-old fix is a proper scoring rule, and this framework's scorecard doesn't carry one yet.

≈ 13 min read

The Prior Was Doing More Than You Thought

A second estimation path made a hidden dependency visible: the media priors were quietly offsetting confounding bias. A penalty chosen by cross-validation cannot do that job, because under confounding the coefficient that predicts best is the confounded one.

From Reading to Measuring

The framework puts these ideas to work: causal identification, calibrated Bayesian MMMs, and experiment design in one loop.