Stop Validating Your MMM With Holdout Error

Every marketing mix model bake-off ends the same way. Three vendors, three decks, and one number that decides it: out-of-sample accuracy on a held-out window. The 4.2% model beats the 5.1% model, procurement has something defensible to write down, and the contract is signed. That criterion measures a quantity nobody is buying. A holdout window scores the model's sum, how close its predicted KPI comes to the actual KPI. Every decision anyone takes from an MMM depends on the split: how much of that same predicted KPI is assigned to media rather than to price, seasonality, distribution, or demand. The sum is disciplined by the data. The split is the free parameter, and two models can agree to four decimal places on the first while disagreeing by a factor of four on the second. No careful analyst can manage around that. Google's own Meridian documentation disowns the criterion in as many words, and the direction of the failure is worse than indifference: the model that predicts best is often the one that attributes worst, because the fastest way to improve a fit is to hand the model a variable that advertising caused. Below, across a grid of 144 defensible specifications on a world whose truth I planted, holdout error and estimand error come out all but unrelated. The winner by holdout MAPE understates working media by 77%.

The Criterion Everyone Competes On

The holdout number acquired its authority for good reasons. No analyst judgment enters the arithmetic, and the whole thing costs one refit and one subtraction. It is also comparable across vendors who share nothing else (different priors, different transforms, different software), because all three can be pointed at the same 26 weeks. And it carries the one statistical lesson every stakeholder has absorbed, that fitting your own data is not evidence. Against a market whose alternative offer is "trust our methodology," a MAPE on unseen weeks looks like the only honest thing in the room.

It is also the wrong tool, and the vendor best placed to profit from it says so. Google's Meridian documentation states plainly that "the goal in marketing mix modeling (MMM) is causal inference, and not necessarily to minimize out-of-sample prediction metrics," and that such metrics are "useful as a preliminary check to make sure the model structure is appropriate and not extremely overparameterized." Its causal-inference page goes further: "A model with 99% out-of-sample R-squared can still be a poor model for causal inference." The model-fit page is more uncomfortable still for a procurement scorecard. It concludes that "it can be safer to have a model that includes all confounding variables and allows enough flexibility in the model structure to get unbiased, causal estimates (such as ROI), even if this means the model is overfit."

Read that last sentence again with a bake-off in mind: the vendor is saying that a model scoring worse on the buyer's criterion may be the one the buyer should want. That admission is more than a vendor's graceful concession. It is the standard position of the statistics literature, arriving late in an applied field. Shmueli (2010) drew the distinction the field never internalized. Explanatory and predictive modeling optimize different objectives, keep different variables, and are validated differently. The same paper showed that an underspecified model can achieve lower prediction error than the correctly specified one, trading a little bias for a lot of variance. Hernán, Hsu and Healy (2019) made the taxonomy operational, separating description, prediction and counterfactual prediction into three tasks with three validation regimes and warning that the third cannot be validated with the tools of the second. Machine learning hit the same wall from the other side. Curth and van der Schaar (2023) set out to settle how to select among models for heterogeneous treatment effects and reported that "due to the absence of counterfactual information in practice, it is usually not possible to rely on standard validation metrics" for the job. Their contribution is deliberately a map of when each proxy misleads rather than a winner. If a field with randomized benchmarks and every incentive to publish a leaderboard cannot produce one for a causal estimand, a bake-off scored on a KPI holdout is not doing something more rigorous. It is doing something easier under the same name. An MMM is a counterfactual-prediction instrument wearing a prediction model's clothes, and we buy it on the clothes.

Definition: the estimand, and what the holdout scores

Write the model's fitted mean for week \( t \) as a baseline plus a sum of channel contributions,

$$ \hat{\mu}_t \;=\; \hat{b}_t \;+\; \sum_{c} \hat{m}_{ct}. $$

A holdout score (MAPE, RMSE, R², expected log predictive density) is a function of \( \hat{\mu}_t \) and \( y_t \) only. It never sees \( \hat{b}_t \) or \( \hat{m}_{ct} \) separately. The estimand you are buying is a function of the parts: the incremental contribution \( \sum_t \hat{m}_{ct} \), the return \( \sum_t \hat{m}_{ct} / \sum_t s_{ct} \), the marginal return under a budget change. Any two models agreeing on \( \hat{\mu} \) score identically and can disagree arbitrarily on every quantity in that second list, subject only to their parts summing to the same total. The holdout tests a different function of the same fit. Evidence about the estimand never enters the calculation.

The Decomposition Is the Free Parameter

The identity above is trivially true, so the interesting question is quantitative: how much room does it leave? In a saturated MMM with a flexible baseline, a great deal, because the baseline and the media block compete for the same variation. Nobody randomizes media spend. It is planned against the same seasonal calendar, promotional calendar and demand forecast that drive the KPI, so the media and baseline regressors are close to collinear by construction. Whatever the fit assigns to one it can nearly un-assign from the other while leaving \( \hat{\mu} \) alone. The holdout prices the total and is silent on the transfer.

The figure below makes the room concrete. Its world is a three-year weekly panel with two channels (a flighted brand channel, and a search channel whose budget chases demand) plus a slow-moving latent demand index that raises sales and pulls search spend upward. Media truly contributes 16.6% of the KPI, and because the truth is planted, every specification's error is knowable. I then fit 144 specifications a competent analyst could defend in a review: four control sets, three baseline shapes, four carryover assumptions, three saturation strengths, each fitted on 130 weeks and scored on the 26 held out.

144 defensible specifications, one planted truth

The upper panel is the weekly media contribution: the planted truth (shaded), the pre-registered causal specification, and whichever specification wins on holdout MAPE. The lower panel is the whole grid, holdout MAPE against the error in total media contribution, colored by what each specification does about the demand confounder. Drag the demand-chasing dial to change how hard the search budget follows demand.

0.60
MAPE, pre-registered
MAPE, holdout winner
Media share: truth
Media share: winner
r (MAPE, |error|)
Error span within 0.5 pp

At the defaults the grid's holdout MAPE runs from 2.7% to 14.3% while the error in total media contribution runs from −100% to +187%, and the correlation between them is r = 0.34. That is weak, and mostly carried by the handful of specifications that fit badly. Where it matters, it vanishes: among the 45 specifications whose holdout MAPE sits within half a point of the best, the estimated media contribution spans −99% to +54% of the truth. The winner on MAPE (2.73%) puts media at 3.8% of the KPI against a truth of 16.6%, while the pre-registered causal specification scores 3.29% on the same criterion and lands within 3.5% of the truth. Ranking by holdout error would have rejected the right answer for one that erases three-quarters of working media.

Three features of that grid deserve attention. The first is the spread within the noise band. Half a point of MAPE is inside the range any reviewer would call a tie (roughly the week-to-week wobble you get from moving the holdout boundary), and inside it the estimand moves by 150 points. The holdout has no resolving power precisely where the decision is being made.

The second is that the failure is not an artifact of the one obviously improper specification family. Drop every model using a downstream variable as a control (the button above does exactly that) and among the 36 survivors within half a point of the best, the media contribution still spans −75% to +72% of truth. Third, the direction is unstable: re-drawing the panel under five seeds, the holdout winner's error came out at −77%, −28%, −26%, −22% and +46%. The criterion does not reliably over- or under-state media. It does not track the quantity at all, which is worse, because a consistent bias could at least be corrected.

None of this should surprise anyone who has read the identification literature. Dew, Padilla and Shchetkina (2024) show that whole families of MMM (a nonlinear response with a fixed coefficient, and a linear response with a drifting one) are often not distinguishable from standard marketing mix data, and that the trigger is autocorrelation in the media variables, which adstock manufactures on purpose. Two such models produce near-identical predictive distributions and opposite budget recommendations. That case has its own post. The consequence here is that no predictive criterion can separate models that are predictively equivalent by construction.

⚠️ A holdout MMM test is not even a forecasting test

Almost every MMM "holdout" hands the model the actual values of every control and media variable in the held-out window and asks it to predict the KPI. Calling that a forecast is generous. It is an interpolation with the right-hand side given. A genuine out-of-time forecast would require forecasting price, distribution, competitor activity and the media plan first, and would score far worse. The number on the deck is already the most flattering version of a criterion that does not measure the thing being purchased. It also explains a quirk in the figure. Specifications with a post-treatment variable win partly because they are handed a column that already contains the answer.

When Predicting Better Means Attributing Worse

Indifference would be tolerable. What makes the criterion actively dangerous is that the two objectives are, in the specific structure of an MMM, often opposed. Two mechanisms do most of the damage, and both are routine practice.

The first is the post-treatment control. Branded search volume, site sessions and email opens are magnificent predictors of sales. They sit in the causal chain immediately upstream of the outcome and immediately downstream of advertising. Adding one improves almost any fit, and transfers to it most of the effect advertising had through it, which is the mechanism advertising works by. This is the Table 2 fallacy in media mix models in its most expensive form. A media coefficient adjusted for a mediator answers a question nobody asked and nobody would fund: "what would this channel do if it could not move branded search?" The analyst who adds it is rewarded on the holdout and punished on the estimand, and no diagnostic inside the fit tells them apart.

The second is the confounder block, where the arithmetic runs the other way. A control that closes a back-door path (a demand index, a competitor-pressure series, a category-volume proxy) earns its place by removing bias in the media coefficients. It does not have to earn its place predictively, and often cannot: it may be noisy, weakly correlated with the KPI, or nearly collinear with a trend the model already has, so adding it costs degrees of freedom and can raise out-of-sample error. That is exactly the trade Meridian describes, and the everyday version of Shmueli's appendix: the correct model is not the best forecaster. A validation regime that ranks specifications by predictive score does worse than stay silent on confounders. It prunes them out of the model, systematically.

The next figure puts a number on how invisible each mistake is. It injects one fault at a time into the same world, holds everything else at the pre-registered specification, and records two things: what the fault does to holdout MAPE, and what it does to total media contribution. To keep the ranking from being an accident of one draw, everything is repeated across sixteen independently simulated worlds. The markers are medians and the whiskers are interquartile ranges.

Six ways to be wrong, and what the holdout notices

Each marker is one injected fault, measured against the pre-registered specification on the same data. Horizontal position is the change in holdout MAPE. Vertical position is the typical error in total media contribution. The shaded strip is the ±0.5 pp band inside which no review would call a difference meaningful, so green squares are the faults a holdout gate would catch and red circles are the ones it misses. Whiskers are interquartile ranges over sixteen simulated worlds. Change the holdout length to check that the ranking is not an artifact of one window.

26
Faults flagged at 0.5 pp
Worst estimand error
…its holdout penalty
Placebo's share of media credit

At a 26-week holdout, omitting the demand confounder inflates media contribution by 66% and costs 0.32 pp of MAPE. A placebo channel (a plausible-looking spend series with no true effect) absorbs a median 32% of the media credit in magnitude while moving MAPE by 0.01 pp. The only fault a half-point gate reliably flags is the over-flexible baseline, which is also the only one of the six that is not a causal fault. It moves the estimand 20%, against 66% for the omitted confounder the gate ignores. That is no coincidence, and it is the key to the next section: the holdout is a good detector of exactly one failure mode, and that mode is not a causal one. The ranking holds at 13, 26, 39 and 52 weeks of holdout.

Reconciling This With Our Own Advice

This blog has told you the opposite. In Common Pitfalls in Statistical Modeling, the prescribed remedy for overfitting is to "score models on estimated out-of-sample predictive accuracy rather than fit," with the PSIS-LOO cross-validation of Vehtari, Gelman and Gabry (2017) as the practical instrument, and the summary rule for comparing models is to score on estimated out-of-sample accuracy, never in-sample fit. That advice is right, and it does not conflict with this post. But the boundary between them is load-bearing, and if you take either rule out of its scope you will do damage.

The reconciliation is that the two posts address different targets. Cross-validated predictive accuracy is the correct criterion for the model's predictive content: how much flexibility the baseline should have, whether a spline trend is buying signal or memorizing noise, whether an extra harmonic is real. Those are questions about \( \hat{\mu} \), which is exactly what a predictive score measures. It is the wrong criterion for the causal content (which variables belong in the adjustment set, what the estimand is, whether the media block is identified), because those are questions about the split, and the split is invisible to it.

The figure above is the boundary, drawn empirically. Six faults, and the half-point holdout gate flags exactly one: the over-flexible baseline, the fault that is overfitting. The gate works perfectly on its own failure mode and is blind to the other five. That is a scope statement, not a contradiction. Use the predictive score to choose the model's flexibility, and never to choose its adjustment set. One exception is large enough to name. When budgets follow the calendar, the trend and seasonality basis is the confounder adjustment, so a predictive score should not set that either (see The Baseline Ate Your Media Effect).

💡 Two more scoping notes worth carrying

A predictive score can falsify, and cannot rank. When a model's forecasts are worthless it is broken, and the holdout will tell you so. That is a real, useful floor. What it cannot do is order two models that both clear the floor. Treat it as a pass/fail gate at a pre-declared threshold, never as the leaderboard.

Ordinary LOO is the wrong shape for MMM data anyway. Leave-one-out assumes exchangeable observations. Weekly MMM data is autocorrelated and the media regressors are adstocked, so leaving out a single week leaves most of its information in its neighbors. Bürkner, Gabry and Vehtari (2020) give the leave-future-out construction that respects the time ordering. If you are going to use a predictive score as a floor, use one whose assumptions your data satisfies.

The Checks That Actually Bind

Deleting a criterion without replacing it is worse than useless, because the vacuum gets filled by whoever tells the best story. The replacement is a workflow rather than another scalar: a sequence of targeted checks, each aimed at a specific way the model could be wrong, reported together, in the posture Gabry, Simpson, Vehtari, Betancourt and Gelman (2019) laid out for Bayesian modeling generally. Model checking there is a collection of comparisons designed to be failable, not a number to be maximized. What follows is that list for an MMM, ordered by how much of the failure space each check can see. The ordering is my judgment, informed by the fault catalogue above.

1. Recovery against an external answer key. Fit the pipeline you intend to ship on synthetic worlds whose causal truth you planted, and compare the intervals to the key. This is the only check that sees structural failure, because the data was not generated by your model. The framework ships a catalogue of such worlds in synth/dgp.py with per-channel truths computed the honest way: the noiseless structural mean minus the same mean with that channel's spend set to zero, which is the counterfactual the model's own estimand targets. Scenarios are labelled with what they violate and whether the truth is even representable in the model's hypothesis space, so under-coverage on unobserved_confounding is a finding and under-coverage on clean is a bug.

2. Refutation and placebo tests. Feed the fitted pipeline inputs that cannot carry a real effect and confirm the effect vanishes. Perturb the data in ways that should not matter and confirm the estimate holds. The framework's suite runs four (permuted media, a permuted KPI, a random common cause, and a random data subset), which between them catch the placebo-channel and over-crediting faults the holdout cannot see. They are also the tests most vulnerable to being underpowered, which is why the framework computes an explicit underpowered flag rather than reporting a bare pass: a "stable" verdict from an estimate that could not have moved is not evidence.

3. Calibration of the inference itself: coverage and SBC. These answer a narrower question than the first two, but they answer it decisively, and they are the reason many teams' intervals are fiction. Simulation-based calibration (Talts, Betancourt, Simpson, Vehtari and Gelman, 2018) checks that posterior ranks are uniform over datasets drawn from your own prior, a property that holds if and only if every central interval has its nominal coverage averaged over that prior. Recovery coverage fixes every parameter at one truth, simulates and refits many times, and counts how often the interval contains it. The figure below is the shape of that output, and its diagnostic value is that two very different diseases produce two visibly different curves.

Nominal versus empirical coverage, and the two ways it fails

Simulated recovery coverage. Each dot is the share of refits whose central interval contained the fixed truth, with a binomial Monte-Carlo interval on that share. The dashed diagonal is perfect calibration. Two dials drive the figure, and they are the two diseases the framework's recovery_diagnosis separates: a posterior that sits in the wrong place, and one whose intervals are too narrow. Use the presets to see their signatures.

0.00
0.45
120
90% interval covers
bias_z
z_spread
Diagnosis

The opening preset is the single most common cause of the complaint "my 90% interval only contained the truth about half the time": an approximate fit. Intervals 2.2× too narrow but perfectly centered have a true coverage of 54% at the nominal 90%, and 24%, 44% and 62% at the nominal 50%, 80% and 95%. A posterior displaced 1.2 standard deviations by an over-tight prior, with correctly sized intervals, covers 67% at the 90% level but only 27% at the 50%: the same headline failure, a different curve shape, and a different fix. The dots are one batch of simulated refits and scatter around that expected curve, which is its own lesson. Drag the refit count down to a dozen and watch a genuinely 90% interval "measure" anywhere from 70% to 100%. The headline number cannot separate the two diseases. bias_z and z_spread can.

4. Predictive checks aimed at decomposition quantities. Ordinary posterior predictive checks compare replicated to observed KPI and pass essentially any model that fits. They are a check on \( \hat{\mu} \), so they inherit the whole problem. What earns its keep are test statistics the decomposition has to get right and the mean does not: the lag-1 autocorrelation of the residuals, their skewness, the fit in weeks where a single channel went dark. A model that reproduces the KPI's mean and variance while failing on residual autocorrelation is telling you that its temporal structure is wrong, and temporal structure is where adstock lives.

5. Agreement with an experiment that was held out of the fit. This is the only check on the list that can see confounding and misspecification simultaneously, because it brings information the observational data does not contain. It has a strict procedural requirement: the experiment must not be in the calibration. Running a geo lift test, folding it into the likelihood, and then reporting that the model agrees with it is circular. Calibration makes the experiment the prior, and a prior cannot then validate the posterior. Register the model's prediction for the experiment before the readout, and score the interval.

Notice where the holdout sits relative to that list. It sees one of the six faults in the figure, the one that is not causal, and the checks above see the others. The number belongs on the report (it describes the model's forecasting behavior, and clients legitimately ask), but in the section that describes the model, not the section that accepts it.

Deep diveWhy recovery coverage cannot see misspecification, and what it therefore proves

The recovery-coverage loop fixes every free parameter at \( \theta^* \), simulates datasets from the likelihood at \( \theta^* \), and refits the same graph on each. Every dataset it grades was generated by the model being graded, so the check is a statement conditional on the model class being right: within that class, does the machinery deliver intervals at their nominal frequency? It detects an approximate posterior, a broken sampler, a prior so tight the data cannot reach the truth, and an estimand mismatch. It structurally cannot detect a wrong adstock family, a missing confounder, or a drifting coefficient, because none of those exist inside the simulation.

That makes the joint reading informative in a way neither check is alone. Under-covering here means the failure is mechanical and the fix is in the inference. Covering here while missing an external answer key means the machinery is fine and the structure is wrong, a more expensive diagnosis and a completely different remedy. The framework's failure-mode table lists model misspecification as "NOT visible when simulating from the model," and says that catching it "needs external truth."

One caveat the tool carries and users forget: Bayesian coverage is only guaranteed averaged over the prior, so an informative prior will legitimately under-cover at a \( \theta^* \) in its tail. A coverage failure at one point is a prompt to look, not a verdict.

The Corollary We Got Wrong

If predictive skill cannot rank specifications for a causal estimand, it cannot weight them either. That corollary is the harder one to see, because model averaging looks like a principled way to avoid the ranking problem rather than an instance of it. The honest response to specification uncertainty is to declare a set of defensible specifications before the fit, run all of them, and report the spread. The tempting next step is to combine them into one number using weights, and the standard machinery for that is stacking (Yao, Vehtari, Simpson and Gelman, 2018), which chooses the mixture maximizing expected log predictive density on held-out data.

Stacking is excellent at what it does. What it does is find the mixture that forecasts best. Applied to a per-channel ROI, it inherits every failure in this post and adds one of its own. Weights that were merely uninformative about which decomposition is right would be survivable. These weights are adverse. They tilt systematically toward the specifications a causal analyst should trust least, because the specifications that predict best are the ones with the leakiest control blocks. In the grid above, a stacking procedure would have concentrated its weight on models that put media at a quarter of its true contribution.

This framework shipped that error. The spec-curve module (the one whose entire purpose is to resist the garden of forking paths by pre-registering a set of specifications) computed LOO-stacking weights and used them to model-average per-channel ROI. Equal weights over the pre-registered set are now the default. Pre-registration has already asserted that every variant is defensible, so there is no post-hoc predictive ground to promote one of them. The module's own documentation now states the argument:

Stacking, its module docstring now says, "chooses the mixture that maximizes expected predictive utility." It answers "which combination forecasts held-out y best?". The docstring goes on: "A spec curve averages a causal estimand (per-channel ROI), and the two objectives come apart precisely where MMM specs differ… Worse, the direction can invert: a spec that overfits the confounder block often predicts better while being less trustworthy for the causal contrast, so stacking can systematically upweight the specs a causal analyst should trust least."

The predictive weights are still computed and reported, deliberately, as a diagnostic rather than an input: divergence from uniform says predictive fit discriminates between your specifications, which is worth knowing and still not a reason to follow it. I record the history because the failure shows how the mistake propagates. Nobody chose to weight a causal estimand by predictive skill. Well-tested Bayesian machinery, whose own documentation is scrupulous about what it optimizes, was imported and pointed at a quantity it was never built for. That is the same move as the bake-off, one level down. It survived review until someone asked what the weights meant. This post is downstream of that question, not upstream of it, and the fix is days old.

The corollary is really about selection in general, so the related trap deserves a name. Choosing the best of many specifications by any criterion, then reporting that specification's interval as though it were the only one fitted, understates uncertainty and exaggerates the winning estimate. That is the winner's curse, and it applies to a holdout leaderboard exactly as it applies to a p-value leaderboard. A spec curve with equal weights is a refusal to hold the contest at all.

What This Does Not Establish

Several things, and the argument is easy to over-drive.

It does not establish that predictive accuracy is worthless. Predictive accuracy is the right criterion for the parts of the model that are forecasts, and MMMs contain such parts: the model's actual use as a planning forecast, and any structure that is not doing confounder-adjustment work. When spend correlates with the calendar, that excludes the trend and seasonality basis (the baseline is the adjustment set). A model that cannot predict at all is broken and should fail. The claim is about ranking causal estimands, not about the criterion's existence.

The simulations are demonstrations, not measurements of prevalence. They are least-squares fits on a stylized two-channel world with a known baseline, Gaussian noise, one latent confounder and one post-treatment variable. Real panels have more channels, more collinearity and heavier tails, and Bayesian fits with informative priors will damp some of the swings shown. The simulation establishes that the failure exists and is large under conditions that are not adversarial. It says nothing about how often the specification a given team ships lands on the wrong side of it. Relatedly, the post-treatment result is flattered by the convention noted earlier: because MMM holdouts hand the model the true value of every control, a downstream variable uses information a genuine forecast would not have. Its damage to the estimand is real. The size of its predictive advantage here is not.

The ordering of checks in the previous section is a judgment. It reflects the fault catalogue I chose, and a different catalogue would reorder it: a team whose main risk is a broken sampler should promote SBC above refutation. The catalogue omits data errors, geo aggregation artifacts, and every failure that lives in the panel rather than the model.

The alternatives have their own failure modes, and I have not priced them. Synthetic answer keys only test the worlds you thought to plant. Refutation tests are often underpowered. Coverage checks cannot see misspecification, by construction. A held-out experiment is only as good as its own design, and buys one channel at one operating point in one window. Running all five costs meaningfully more than computing a MAPE.

The Acceptance Gate

If a holdout number should not decide a bake-off, something has to. Here is what I would write into a contract, in the order the work happens. None of it requires methodology the field does not have, and all of it requires the vendor to commit before seeing results, which is the part that does the work.

Before the fit. The estimand, in one sentence, with units: incremental KPI per dollar over a stated window, or per thousand impressions where spend is not the modeled variable. A directed acyclic graph or an explicit list of the adjustment set with a stated role for every control, so that a post-treatment variable has to be argued for rather than slipped in. And a pre-registered set of defensible specifications, filed with a timestamp, together with the rule for how disagreement across them will be reported.

At delivery. Coverage evidence, not assertion: a recovery-coverage run at a fixed truth reporting empirical coverage with Monte-Carlo bounds at 50, 80, 90 and 95%, and an SBC run for the same specification. A refutation suite with its power flag shown, so passes from tests that could not have moved are not counted. The spec curve across the pre-registered set with equal weights, its per-channel range stated as prominently as the point estimate. And a tier on every headline number saying whether it is experiment-validated, model-identified, or prior-dominated, because a number whose posterior barely moved off its prior is the analyst's assumption wearing a model.

After delivery. One experiment, designed against the model's own uncertainty, held out of the calibration, with the model's prediction for it registered before the readout. That is the only clause on this list that can indict the whole apparatus, and it is the one worth paying for.

And keep the holdout. Report it, with the honest caveat that the controls were given, as a description of the model's forecasting behavior and as a pass/fail floor against a threshold declared in advance. What it must not be is the thing that decides. The number that decides should be the one you are buying: how much of the KPI media actually caused, with an interval you have earned the right to quote.

Takeaways

  • A holdout score is a function of the fitted mean \( \hat{\mu}_t \) alone, while every MMM estimand is a function of how \( \hat{\mu}_t \) splits into baseline and per-channel contributions. The criterion and the deliverable are different functions of the same fit.
  • Across 144 defensible specifications on a planted truth, holdout MAPE and media-contribution error correlate at r = 0.34, and inside the half-point band no reviewer would call meaningful, the estimand spans −99% to +54% of truth. The MAPE winner put media at 3.8% of the KPI against a truth of 16.6%.
  • The relationship is often inverted, not merely absent: a post-treatment control improves the fit and steals the effect it mediates, while a genuine confounder can cost predictive accuracy and remove bias. Meridian's documentation says the safer model may be the overfit one.
  • This does not retract the standing advice to score on estimated out-of-sample accuracy. That rule is right for a predictive target. In the fault catalogue here, a half-point holdout gate flags exactly one of six faults: the over-flexible baseline, which is overfitting. Use the predictive score to set flexibility, never to choose the adjustment set, and note that a calendar-correlated baseline is part of the adjustment set.
  • Replace it, ranked by failure space covered: recovery against an external answer key, refutation and placebo tests, coverage and SBC, predictive checks aimed at decomposition quantities, and agreement with an experiment that was held out of the calibration.
  • Predictive weights are not valid weights for averaging a causal estimand. Stacking optimizes forecast utility and will upweight the leakiest specifications. This framework shipped that error in its spec-curve module and now defaults to equal weights over the pre-registered set, reporting the stacking weights only as a diagnostic.

References