Modern Measurement Research

Research essays on how marketing measurement actually works — and where it quietly fails. Written from the primary literature: causal inference, geo experiments, Bayesian media mix modeling, and the design of the next experiment.

Part I

Why Measurement Is Hard

Before any method can help, it pays to understand why the naive answers are wrong. Observational data flatters advertising, and untangling cause from correlation takes deliberate machinery.

≈ 9 min read

Activity Bias: Why Observational Ad Measurement Flatters Itself

People who browse more see more ads and do more of everything else online. That single fact can inflate observational estimates of ad effectiveness by orders of magnitude.

≈ 15 min read

Uncovering Causal Estimates from Non-Experimental Data

A working tour of the identification strategies that stand between you and a causal claim when you cannot randomize: adjustment, instruments, panels, discontinuities, and sensitivity analysis.

≈ 13 min read

The Table 2 Fallacy: When a Coefficient Isn't an Effect

One regression fits one causal question, but the output table tempts you to read every coefficient as an effect. Only the exposure's is identified — the rest mix confounding, mediation, and precise nonsense.

≈ 14 min read

Attribution Is Not Incrementality

Attribution splits credit for conversions that happened; incrementality asks which would not have happened without the ad. The gap is largest exactly where budgets concentrate — retargeting and branded search.

Part II

Quasi-Experimental Methods

The workhorse designs for measuring incrementality without a full randomized experiment — and the assumptions each one quietly leans on.

≈ 14 min read

Geo Experiments: TBR, Power, and the Stationarity You're Assuming

Time-based regression turns a geo split into a causal estimate of ad-driven lift. It works — right up until the pre-period relationship between markets stops holding.

≈ 13 min read

Switchback Experiments: Randomizing Time, Not Users

When a marketplace makes user-level A/B tests biased, randomize the whole system on and off over time instead. Carryover sets the switching interval — and long adstock is what breaks it.

≈ 15 min read

Synthetic Control, Done Right

When only one market gets the treatment, a weighted blend of untreated markets can stand in for the counterfactual — if you respect the method's bias theory, placebo inference, and requirements.

≈ 13 min read

Modern Staggered Difference-in-Differences

Two-way fixed effects breaks when treatment rolls out in waves and effects differ across markets or build over time. Group-time average treatment effects rebuild DiD from clean comparisons — and tell you exactly what you're averaging.

≈ 14 min read

Pretest With Caution: Testing for Parallel Trends

Checking that pre-treatment leads are insignificant feels like proof of parallel trends. It isn't: the test is underpowered against the trends that bias you most, and conditioning on passing it distorts the estimate itself.

≈ 13 min read

CausalImpact: Counterfactuals from Bayesian Structural Time Series

Fit a state-space model to the pre-period, forecast the world where the campaign never launched, and read the lift as the gap — with honest uncertainty that widens over time.

Part III

Models and Their Failure Modes

Media mix models encode strong assumptions about carryover, saturation, and selection. These essays cover the modeling choices that matter and the mistakes that survive peer review.

≈ 14 min read

Carryover and Shape Effects in Bayesian MMM

Adstock and saturation are where a media mix model's causal claims live. How the carryover and Hill-shape transforms work, why their parameters are hard to identify, and what priors buy you.

≈ 13 min read

The Table 2 Fallacy in Media Mix Models

An MMM's control coefficients are adjustment terms, not effects — and a downstream "control" like branded search quietly strips an upper-funnel channel's credit. Good and bad controls in a media mix model.

≈ 14 min read

Measuring the Long Run: Surrogate Outcomes in Media

Budgets need long-run effects; experiments observe short-term proxies. A single surrogate rarely satisfies Prentice's criterion — but an index of many short-term signals, validated against occasional truth, can bridge the gap.

≈ 12 min read

The Experiment Is the Prior: Calibrating a Media Mix Model

An observational MMM coefficient is not causal — spend is endogenous and collinear. A geo lift test identifies what the likelihood cannot, and calibration carries it back in as a prior.

≈ 13 min read

Adstock Dynamics and the Timing of Sequential Experiments

Carryover doesn't just complicate models — it contaminates experiments. How long to wash out, how long to measure, and when a lagged effect makes back-to-back tests read each other's answers.

≈ 12 min read

Common Pitfalls in Statistical Modeling

A field guide to the failure modes that quietly invalidate applied statistical work — from confounded regressions and the R-squared trap to the garden of forking paths.

≈ 14 min read

Hundreds of Models, One Winner: Multiple Comparisons and the Winner's Curse

Select the best of hundreds of model runs and its performance is almost surely exaggerated. Type S/M errors, the winner's curse, and why partial pooling beats correction-by-penalty.

≈ 13 min read

P-Hacking and the Thin Evidence of a P-Value

A significant result is both weak evidence and easy to manufacture. Researcher degrees of freedom inflate false positives — and a p just under 0.05 leaves the null with about a one-in-four chance of being true.

≈ 15 min read

Does Your Model Work Before It Sees Data?

Prior predictive checks catch priors that imply impossible outcomes; simulation-based calibration verifies the posterior computation itself — two checks that run before a single real observation.

≈ 15 min read

When Sampling Fails

R-hat, ESS, and divergences each point to a different disease — and only one is cured by running longer. An escalation ladder for failed Bayesian fits, ending with SMC as the multimodality second opinion.

Part IV

Designing the Next Experiment

Measurement is a loop, not a report. Information theory tells you which experiment to run next; bandits tell you how to earn while you learn.

≈ 15 min read

From Lindley to Deep Adaptive Design: 65 Years of Bayesian Experimental Design

The expected-information-gain criterion is older than most measurement teams realize. From Lindley's 1956 measure to variational estimators and amortized design policies.

≈ 9 min read

A Geo-Holdout Is a Bayesian Experimental Design: Computing Its EIG

Encode a holdout test as a design vector, treat the MMM posterior as the prior, and the value of the experiment becomes a computable quantity — expected information gain.

≈ 13 min read

BED vs. Bayesian Optimization vs. Bandits for Media Experimentation

Three adaptive frameworks, three different goals: learn the model, find the optimum, or earn while learning. Choosing the wrong one wastes budget in predictable ways.

≈ 14 min read

Thompson Sampling: A Practical Tour

Sample from the posterior, act as if the sample were true, repeat. Why this one-line algorithm earns near-optimal regret, and what breaks when feedback is delayed or non-stationary.

≈ 13 min read

Continuous Learning in Media Measurement (with Interaction Effects)

Closing the loop: a standing program of designed experiments that keeps a response model current — including the channel-interaction effects one-off tests never identify.

From Reading to Measuring

The framework puts these ideas to work: causal identification, calibrated Bayesian MMMs, and experiment design in one loop.