How should a geo-lift or A/B experiment result become a prior in a Bayesian media mix model, and what should happen when the experiment and the MMM disagree?

Summary

An experiment does not estimate an MMM parameter. It estimates one functional of the response curve — incremental sales per incremental dollar for one channel, between two specific spend levels, in specific geos, over one window including carryover — whereas the MMM’s unknowns define the whole curve. So the clean way in is as an extra likelihood term on the MMM-implied version of the same counterfactual, , not as a prior pasted onto ; with several experiments, as a hierarchical layer whose between-experiment spread is the honest prior width. When the two disagree, the normal-normal update will silently average them into an answer neither supports, so the protocol is: overlay prior and likelihood on the functional, align estimands, audit both sides, then choose explicitly who yields — heavy tails (the model yields to the data it trusts more), an explicit bias term on the observational side (the MMM yields to the experiment), or reported conflict plus a new experiment designed to resolve it.

Answer

1. Why this matters: the MMM is prior-dominated

Bayesian Estimation and Priors for MMM is blunt: with a couple of years of weekly data “the posterior may look almost the same as the prior.” In Jin et al.’s simulations the Hill curves are underestimated by 18–33% at two years of data and “there is no universally ‘correct’ prior”; in the shampoo application the retention-rate posterior ”≈ its prior” and the optimal-mix posterior is bimodal (MMM Model Selection and Application). The Bayesian framing was adopted precisely “to incorporate prior knowledge.” On the other side, Activity Bias in Advertising shows observational ad estimates can be wrong by two orders of magnitude (872% vs a randomised 5.4%) and that “more data doesn’t help … this is a bias problem, not a variance problem” (Observational vs Experimental Methods in Advertising). An experiment is therefore both the most valuable information an MMM can receive and the thing most likely to contradict it. Geo-Experiment Methodology - Overview names the use directly: a geo iROAS “is exactly the kind of ground truth used to calibrate or validate an MMM’s channel coefficients.”

2. Two estimands that only look alike

ExperimentMMM
ObjectOne number with an SE: or (Time-Based Regression Estimator for Geo Experiments)A function:
Spend contrastFrom actual baseline to perturbed (go-dark, heavy-up) in the treated geosAny; ROAS zeroes the channel, mROAS adds 1% (ROAS, mROAS, and Optimal Media Mix)
TimeIntervention + cooldown weeks of one seasonAverage over the whole fitting window, carryover summed to
PopulationTreated geos (volume-weighted for TBR; average geo on a log scale for SDID)National or all-geo
IdentificationRandomisation or a panel counterfactualRegression on observational spend variation

Consequences:

  • A go-dark holdout matches the MMM’s zero-out ROAS over that window. A heavy-up measures a finite difference above current spend, which on a concave curve is below average ROAS and approaches mROAS only for small perturbations. Comparing a heavy-up iROAS with the MMM’s average ROAS manufactures a “disagreement” that is just curvature. TBR’s design note flags the same thing: larger spend intensity “risks hitting diminishing marginal returns.”
  • One experiment pins one point-to-point secant, not the curve. Shape (Saturation) Effects shows very different triples give nearly identical curves in range: the curve is estimable, the parameters are not. An experiment therefore constrains a ridge in parameter space; separating slope from saturation needs a second experiment at a different spend level.
  • Carryover must be on both sides. MMM ROAS includes the post-change period; an experiment read without cooldown understates lift (Q - Carryover Dynamics and the Timing of Sequential Media Experiments).

The mapping is then a definition, computed per posterior draw (never from posterior means — ROAS, mROAS, and Optimal Media Mix):

with the experiment’s own geos, weeks, spend paths and outcome definition. This is Prior Distributions’ advice made concrete: “elicit on the predictive scale, not the parameter scale.”

3. Three ways to let the experiment in

A. Experiment as likelihood (default). Add one line per experiment:

In the taxonomy of Modeled and Unmodeled Data, is modeled data, and the spend paths are unmodeled data. It is mathematically a prior on the functional — “data models sometimes become priors” — but it respects the nonlinearity, composes across experiments, and needs no re-parameterisation. The geo-holdout Q&A already treats the MMM as the likelihood for a design ; this is that update done with a sufficient summary instead of raw geo-weeks. SDID for Geo Experiments and Marketing Panels makes the same point from the estimator side: a point estimate with a Gaussian interval “can still be used as a likelihood summary for MMM calibration,” while BSTS or TBR hand over a full posterior.

Double counting ( synthesis)

If the MMM is fit on geo-level data that include the test geos and weeks, the experiment’s sales are already in the likelihood. Either keep the raw data and drop the summary term (the randomised spend shock is then exogenous variation the MMM sees directly), or keep the summary and mask those cells. For a national MMM the overlap is usually negligible.

B. Experiment as a prior on a parameter (only with care). A half-normal on cannot encode an iROAS: the same ROAS corresponds to different as move. If the tooling only accepts parameter priors, tune the prior and verify with a prior predictive check: simulate from the prior, compute , and confirm its distribution matches the experiment’s. This is a level-5 “specific informative” prior, so follow the documentation rule — “write a sentence about each parameter” tracing every number to its experiment.

C. Several experiments: a hierarchical or meta-analytic prior. “The prior for any given study represents the distribution of average treatment effects among a hypothetical population of problems” (Constructing Priors for Effect Sizes). With experiments on a channel (different quarters, creatives, regions):

is transport variance — how far a true effect in one window sits from the curve’s average. It is what stops a single tight experiment from over-ruling two years of data about a different period. Hierarchical Models warns that with the posterior for collapses toward zero unless given a half-Cauchy or half- hyperprior with a domain-informed scale; with many experiments across brands, empirical Bayes estimation of the population is the cheap version. Jin et al. recommend the same move — pool brands or geos “to manufacture more informative priors” rather than lengthen history, because market conditions drift.

What the update does (normal-normal arithmetic as in Constructing Priors for Effect Sizes)

MMM-only posterior for the go-dark ROAS of paid social: . Geo test: . Precision weighting gives

Add transport sd to the experiment (): the mean moves to , sd . Both answers lie in a region neither source favoured — which is why §5 exists.

4. How wide should be?

The reported standard error is a lower bound. Add, in variance:

  1. Estimator uncertainty. The same geo panel gives different lifts under TBR, CausalImpact, SC and SDID (Q - Comparing Geo-Test Estimators from TBR to Synthetic DiD); TBR’s interval is valid only under its stability assumption (TBR Design Sensitivity and the Stationarity Assumption).
  2. Transport (): time, creative, geo mix, spend level.
  3. Type M inflation. “Statistically significant results from underpowered studies are therefore likely to be exaggerated” (Type S and Type M Errors). If the test is in the deck because it was significant, shrink it first: the early-childhood example turns into under a literature prior. Feeding the raw likelihood (route A) does this automatically, provided the MMM’s own prior on the functional is realistic rather than flat.

A noisy experiment then contributes little, by arithmetic rather than by judgement call.

5. When they disagree: a protocol

The normal-normal model hides the conflict ( Tail Behavior and Prior-Likelihood Conflict)

With against a prior of equal information, the posterior is — “contradicting both prior and likelihood” — and “computation remains smooth”: no divergences, no warning. “Plotting prior and likelihood together is the only check that catches it.”

Step 0 — Detect. For each experiment overlay (i) the MMM-only posterior of and (ii) the experiment likelihood. Run power-scaling sensitivity (Influence of Likelihood and Prior): a functional sensitive to both prior and likelihood scaling signals conflict; trust the importance-sampling answer only when Pareto .

Step 1 — Align estimands (§2): same window with cooldown, same spend contrast (secant vs. average vs. marginal), same geos and outcome. Many disagreements end here.

Step 2 — Audit the experiment. Estimator multiverse, placebo/A-A backtests, spillover into control geos, power and Type M, whether geos were randomised or hand-picked.

Step 3 — Audit the MMM. Is the functional prior-dominated (posterior ≈ prior)? Residual autocorrelation out to lag ~15 weeks was a misspecification signal in the shampoo model. Is spend correlated with unmodelled demand, promotions or other channels — the aggregate analogue of activity bias (synthesis: budgets follow expected demand, so the bias is typically upward for demand-capturing channels)?

Step 4 — Choose who yields, explicitly.

Remedy (Gelman et al.’s table)MMM translationConflict resolves towardUse when
Thick-tailed termStudent-/Cauchy on the experiment termthe MMM’s time-series datathe experiment’s transportability or execution is doubtful
Explicit bias term, ; experiment informs , budget decisions use the experimentthe experiment is clean and the MMM is observational — the usual case
Thin tails + diagnosticsGaussian terms, conflict reportedneitherstakeholders should see both numbers

The bias-term row is the fix Relating a Model to Subject-Matter Assumptions prescribes when estimates come “from observational studies rather than controlled experiments,” and it is Plausible GMM’s move: replace the dogmatic assumption that the exogeneity moment holds exactly () with a proper prior over the violation. Two of its lessons carry over. First, no free lunch — admitting possible bias “necessarily yields wider, less precise inference.” Second, and “are not jointly identified,” so from observational data alone the prior on never washes out; the experiment is what identifies the bias, and a channel never tested keeps its full bias uncertainty. Expect the counter-intuitive behaviour Gelman et al. document: is non-monotonic — the larger the gap, the more is attributed to bias and the less the observational data move the causal estimate.

Step 5 — If still unresolved, make it a design problem. Choose the next test to maximise targeted EIG on the contested functional, ideally at a different spend level so curvature is identified (Q - Encoding a Geo-Holdout as a Bayesian Experimental Design and Computing Its EIG).

Step 6 — Document. For each experiment term: source, dates, geos, estimator, raw SE, inflation applied, window, and the conflict plot.

Practical Implications

  • Default recipe: one likelihood term per experiment on a matched counterfactual ; hierarchical per channel; a per-channel observational bias term with a tight-centred, heavy-tailed prior; decisions (mROAS, optimal mix) computed from the causal curve, per draw.
  • Match experiment type to MMM functional: go-dark ↔ ROAS; small heavy-up ↔ mROAS; large heavy-up ↔ a secant you must compute explicitly.
  • Plan experiments in pairs of spend levels if you want them to discipline saturation and not only scale.
  • Never trust a calibrated MMM without the overlay plot; smooth sampling is not evidence of agreement.
  • Age your experiments (synthesis): inflate or with time since the test — the informal version of the “hierarchical time series model” that Constructing Priors for Effect Sizes says historical priors approximate.
  • For agent-based models the same device applies: a lift test becomes one more target statistic whose simulated counterpart is the ABM’s geo-holdout counterfactual (Q - Using SMM to Calibrate Agent Based Models).

Source Notes

NoteRelevance
Bayesian Estimation and Priors for MMMPrior specifications, prior dominance, sensitivity of Hill to the prior
ROAS, mROAS, and Optimal Media MixCounterfactual definitions of ROAS/mROAS; plug in draws, not means; optimal-mix variance
Shape (Saturation) Effects · Carryover (Adstock) Functional FormsCurve identifiable, parameters not; carryover window
MMM Model Selection and Application · Bayesian Media Mix Modeling - OverviewSmall-sample bias numbers, residual autocorrelation, pooling for priors
Constructing Priors for Effect SizesPrior as population of effects; meta-analytic priors; early-childhood shrinkage example
Tail Behavior and Prior-Likelihood ConflictHidden conflict, Cauchy vs bias-term remedies, non-monotonic posterior mean
Influence of Likelihood and PriorPower-scaling sensitivity and the conflict diagnostic
Prior Distributions · Prior Predictive Checking · Modeled and Unmodeled DataPredictive-scale elicitation, informativity ladder, documentation, variable taxonomy
Relating a Model to Subject-Matter AssumptionsBias term when estimates are observational
Hierarchical Models · Empirical Bayes - OverviewPartial pooling across experiments; hyperprior on with few groups
Time-Based Regression Estimator for Geo Experiments · TBR Design Sensitivity and the Stationarity Assumption · SDID for Geo Experiments and Marketing PanelsWhat the experiment delivers (iROAS posterior / Gaussian summary) and its validity conditions
Activity Bias in Advertising · Observational vs Experimental Methods in AdvertisingSize and nature of observational bias
Type S and Type M ErrorsExaggeration of significant, underpowered results
Plausible GMM - OverviewPriors over misspecification; non-identification of bias; no free lunch
Q - Encoding a Geo-Holdout as a Bayesian Experimental Design and Computing Its EIGMMM as likelihood for a design; choosing the next experiment

Gaps

  • No ingested source on MMM calibration practice (ROI-parameterised priors, lift-calibrated loss terms, experiment-based bias correction). Sections 3–5 are synthesis from the MMM, workflow and geo-experiment notes.
  • No coverage of power priors or commensurate priors for discounting historical experiments; the ageing rule is improvised.
  • User-level A/B lift → MMM units (conversions per exposed user to sales per dollar, platform-attributed outcomes, reach scaling) is not covered; this answer is solid only for geo-level tests.
  • Time-varying MMM coefficients are absent, so “the experiment measured a different period” can only be absorbed by .
  • Plausible GMM - Overview lacks the formal §4 results, so the mapping to an MMM bias term is by analogy.

Follow-Up Questions

  • How should the MMM be re-parameterised so that channel ROAS at reference spend is itself a parameter with a prior?
  • What prior scale for the observational bias is defensible per channel type (search vs. TV)?
  • Given the current posterior, which spend level for the next geo test most reduces uncertainty about ?
  • How fast should an experiment’s weight decay, and can that be estimated from repeated tests on one channel?