The Baseline Ate Your Media Effect

Every marketing mix model contains a term whose only job is to absorb demand you cannot measure: a growth rate, ten Prophet changepoints, a twenty-knot spline, a Fourier stack, a Gaussian process, a local level. It carries no business meaning, appears on no slide, and is usually set by whoever last noticed the fit improve. It is also, in every model whose budgets follow the calendar, the adjustment set standing between the media coefficients and the confounding nothing else can remove. Google's Meridian documentation says so in its own words: “If time is an important confounder between media and the KPI, then the bias-variance trade-off in estimating time effects translates to a bias-variance trade-off in estimating the causal effects of media.” Turn that dial and the ROI moves. What makes it dangerous is the direction of the error signal. On a planted world below, a baseline stiff enough to look safe puts the demand-chasing search channel at 35.0% of the KPI against a truth of 13.2%, and it reports the narrowest interval in the sweep, 6.5 points wide, containing the truth in 0 of 200 replications. Loosening it walks the estimate onto the truth and widens the interval to 15 points. Precision and accuracy move in opposite directions, so nothing in the output warns you, and cross-validated error is minimized at a third setting again. Baseline flexibility is a causal hyperparameter wearing a fit hyperparameter's clothes. It cannot be tuned. It has to be swept.

Time Is the Confounder You Adjust With

The structural claim is short. Budgets are set against the calendar, and the calendar is no randomizer. Retail spends into Q4 because Q4 sells, insurers buy open enrollment, and everyone's annual plan is anchored on last year's shape. The KPI moves on the same calendar for reasons that have nothing to do with advertising: weather, category seasonality, holiday timing, the macro cycle. Calendar time is therefore a common cause of spend and of sales, and the part of it you did not measure is a textbook back-door path from media to the KPI. Most MMMs have exactly one instrument for blocking that path, a flexible function of time. It is a confounder adjustment built out of the index column, and it asks for no data you do not already have.

None of that is an outside critique of anybody's software. It is the position taken by the documentation of the most widely deployed open MMM in the market. Meridian's page on setting knots states the trade-off in exactly the terms a causal analyst would use: “More knots reduce bias in time effect estimates, while fewer knots reduce variance in time effect estimates”. The page then transmits that trade-off, as quoted above, into the media effects themselves, and documents the knot count as an analyst decision about where on a bias-variance frontier to sit. Almost nobody treats it that way. It gets set to the default, or nudged until the fit stops looking lumpy, and it never appears in the assumption log.

Two neighboring arguments locate this one. The Table 2 Fallacy in Media Mix Models establishes that MMM control coefficients carry no clean causal reading, and notes in passing that over-shrinking a confounder “re-opens the back-door path into the channel estimates”. The same mechanism, applied to columns someone named. The new subject is the term that is not in the coefficient table at all, a basis nobody interprets, which therefore escapes every review that would catch a suspiciously tight prior on price. And Your Demand Proxy Is Not a Control Variable shows that a measured stand-in for latent demand does not buy partial de-confounding by partial adjustment. The time basis is the other way of proxying the same latent quantity: unmeasured, fit from the residuals. It needs no external data and cannot be contaminated by advertising the way a branded-search series can. What it can represent is demand that moves as a function of calendar time common to the whole panel, and nothing else. A shock in three DMAs is invisible to it. Its flexibility, meanwhile, is chosen by fit.

A Baseline Is a Saturated Regressor Block

To see why the transfer is mechanical rather than incidental, stop thinking of the baseline as a curve and start thinking of it as regressors. Hodges and Reich, writing about spatial models in 2010, showed that adding a spatially correlated random effect to a linear model is equivalent to adding a saturated collection of canonical regressors whose coefficients are shrunk toward zero, with the spatial structure determining both which regressors you get and how hard each is shrunk. Their title is the finding: Adding Spatially-Correlated Errors Can Mess Up the Fixed Effect You Love. Replace “spatial map” with “time index” and you have an MMM's baseline. A Prophet changepoint block, a B-spline basis, a Fourier stack and an HSGP are the same object: a design matrix \( B \) with coefficients \( \gamma \) that a prior pulls toward zero.

Written that way, the bias has a closed form the framework's own README already carries, for the one-confounder case. Let \( C \) be a confounder that drives both spend \( X \) and the KPI, and suppose the fit shrinks \( C \)'s coefficient by a factor \( s \in [0,1] \), so \( \tilde\gamma = s\gamma \). Then

$$ \hat\beta \;\xrightarrow{p}\; \beta \;+\; (1-s)\cdot\gamma\cdot\frac{\mathrm{Cov}(X, C)}{\mathrm{Var}(X)}. $$

Two readings matter. First, only \( s = 1 \), which is to say no shrinkage at all, removes the bias. Shrink a confounder's coefficient by 90% and 90% of the omitted-variable bias survives. Second, the multiplier \( \mathrm{Cov}(X,C)/\mathrm{Var}(X) \) is a property of the world, not of the estimator. Shrinking \( \gamma \) does not shrink the correlation between spend and demand. You cannot regularize your way out of a design problem.

Deep diveThe same formula, one direction at a time

For a basis block the scalar \( s \) becomes a spectrum. Take the singular value decomposition \( B = UDV^\top \). Under a Gaussian prior \( \gamma \sim \mathcal{N}(0, \tau^2 I) \) and noise variance \( \sigma^2 \), the posterior mean of \( \gamma \) is the ridge estimate with penalty \( \lambda = \sigma^2/\tau^2 \), and the \( k \)-th canonical direction is shrunk by

$$ s_k \;=\; \frac{d_k^2}{d_k^2 + \lambda}, \qquad \lambda = \sigma^2/\tau^2 . $$

Directions with large singular values (the basis's smooth, high-energy shapes) are barely shrunk. Wiggly low-energy directions are shrunk almost to nothing. The shrinkage is not uniform, and the basis decides which shapes survive.

Now project the media regressor onto those directions, \( x = \sum_k c_k u_k + x^\perp \). Where \( x \) loads on a single direction \( u_k \), the scalar algebra above applies exactly with \( s = s_k \), giving transferred bias \( (1-s_k)\gamma_k c_k / \mathrm{Var}(x) \). In general \( \beta \) and \( \gamma \) are estimated jointly and the cross-terms do not vanish, so the sum over \( k \) is a first-order reading rather than an identity. The qualitative content survives and is the whole argument: the media coefficient inherits bias in proportion to how much of the channel's variation lies in directions the baseline was not allowed to fit. Two knobs control that: the number of basis functions, which decides which directions exist, and the prior width \( \tau \), which decides how much of each is kept.

So dropping the knot count and tightening the prior are one problem with two spellings, and a spec set that varies knot counts while holding a tight prior fixed has swept the smaller of the two.

The practical version of the spectrum is a statement about timescales. A basis with \( k \) interior knots over \( T \) periods resolves variation down to roughly \( \Delta t \approx T/(k+1) \) and cannot represent anything faster. Three years of weekly data with ten Prophet changepoints resolves about a fourteen-week wave: any demand movement slower than a quarter is absorbed into the baseline, and anything faster passes through into the residual, where the media block is waiting. The same cut applies to media. A channel flighted in three-week bursts keeps almost all of its variation on the fast side of that line and is identified. A channel whose budget steps once a quarter lives on the same side of the line as the baseline, and the two are nearly the same regressor.

Definition: concurvity

Concurvity is the nonparametric generalization of collinearity, introduced by Buja, Hastie and Tibshirani (1989) for additive models: a covariate exhibits concurvity when it is well approximated by a smooth function of the model's other terms. Perfect concurvity makes the parameters unidentifiable, exactly as perfect collinearity does in a linear model. Partial concurvity inflates variance and transmits bias the way a high variance inflation factor does.

For an MMM the question is the concurvity of each channel with the time basis: how much of this channel's spend is a smooth function of the week number? It is a one-line regression: project the channel's modeled variable onto the baseline basis and report \( 1 - R^2 \), the fraction of its variance surviving as identifying variation. Wood (2017) treats a version of this as a standard GAM diagnostic. Almost no MMM report contains it.

Meridian's endpoints make this concrete. Its documentation notes that with knots = 1, “all time periods are measured with a single parameter which is equivalent to saying time has no effect”, and that with knots = n_times, “each time period gets its own parameter and so the effect of a given time period is estimated using only data from that time period.” The defaults split along the line the timescale argument predicts: national models default to one knot, geo models to one per period. In a national panel with one row per week, a parameter per week is a perfect fit that leaves no time-series variation for anything else, and media would not be identified at all. Geo models can afford it because a period effect shared across regions still leaves within-week cross-geo contrasts. Calling the knot count a smoothing preference misses what it does. It decides which comparisons the estimate rests on.

Narrower and Wronger

The figure below plants a world and sweeps the baseline: three years of weekly national data, two channels, one slow latent demand index that raises sales directly. Brand is flighted three weeks on, three weeks off, so it shares almost nothing with demand. Search has its budget re-set every eight weeks to chase where demand is going, which is what a competent performance team does. Search genuinely delivers 13.2% of the KPI, brand 5.3%. Demand is never given to the model. So the only thing standing between the search coefficient and the confounding is the baseline, swept here from a single linear growth rate to a changepoint every three weeks. Watch what the intervals do while the point estimate travels.

The bias curve, and the band that tightens around the wrong answer

Estimated contribution of the demand-chasing search channel, as a share of KPI, against baseline flexibility. The dashed line is the planted truth (13.2%), and the shaded ribbon is the 90% interval. Drag demand-chasing to change how hard the search budget follows demand. At zero the channel is unconfounded and the curve flattens onto the truth. Baseline flexibility runs left to right as a changepoint every 156 weeks (a plain linear trend) down to every 3 weeks. Least-squares fits stand in for the posterior.

0.55
Stiffest baseline
Its 90% width
Loosest baseline
Its 90% width
Truth
Intervals covering truth

At the defaults the stiffest baseline (one linear growth rate over three years) puts search at 35.0% of the KPI against a truth of 13.2%, with the narrowest interval in the sweep (6.5 points wide, [31.7, 38.3]). A changepoint every eight weeks lands on 14.2% with an interval half again as wide. The point estimate walks 21 points while the interval grows steadily the whole way, 6.5 to 15.2 points. Only 4 of the 10 specifications cover the truth, and they are the four least confident ones.

The direction of the interval is the whole problem. Everything an analyst uses to sanity-check a coefficient (interval width, the ratio of estimate to uncertainty, whether zero is excluded) gets better as the baseline gets stiffer and the estimate gets worse. Nor is it a coincidence: the stiff baseline leaves 90% of the search channel's variance orthogonal to the model's other terms, the loose one 6%. The standard error is small in the first case because the regressor looks beautifully separated from everything else. The model is confident because it was never asked the hard question.

Repeating the exercise over independent draws of the same generator turns that into a coverage statement. Across 200 replications, the 90% interval on the search channel's contribution contained the truth in 0 of 200 fits under the linear baseline, with a median width of 6.2 points. Coverage rose to 21% with a changepoint every 16 weeks, to 72% at every 8 weeks, and to 88.5% at every 3 weeks, where the median width had grown to 15.3 points. Mean bias fell monotonically from +188% to +9% along the same axis. Nominal coverage is bought entirely with width, and the setting that a reviewer would praise for its tight intervals is the setting whose intervals are never right. (That replication study is a separate offline computation, not the single planted world in the figure.)

⚠️ Your coverage diagnostic cannot see this

The obvious defense is to check coverage, and this framework ships the tool: diagnostics/coverage.py fixes a parameter vector, simulates datasets, refits, and counts how often each interval contains the truth. It will not catch this failure, and its own documentation says why: “it simulates data FROM the model, so it can never detect real-world misspecification (wrong adstock/saturation family, unobserved confounding, time-varying effects).” Simulate from a model whose baseline is too stiff to represent demand and the generated data contain no demand confounding either. The stiff model then recovers its own truth with nominal coverage and a satisfyingly narrow interval. The module states the disambiguation. Under-coverage there is mechanical: an approximate fit, a broken sampler, priors too tight to move. Coverage there, combined with failure against an external answer key, means the gap is structural. This one is structural, so only the external key finds it.

One Dial, Two Objectives

If the interval will not warn you, the natural next thought is to let the data choose: fit at several flexibilities, keep the one that predicts best. This is where the argument acquires a precedent from a literature that had to settle exactly this question under public scrutiny. Time-series studies of air pollution and mortality regress daily deaths on daily pollution with a smooth function of calendar time absorbing season and long-term trend. That is structurally the same model as an MMM: same confounder, same nuisance basis. The degrees of freedom of that smooth were contested for years, because the estimated pollution effect moved with them.

Peng, Dominici and Louis (2006) settled the methodology in a simulation study, and their conclusions transfer without modification. The amount of smoothing needed for a small-bias estimate of the exposure coefficient, they found, can differ from the amount needed to estimate the nuisance function well. More pointedly, they warned that “the automatic use of criteria such as generalized cross-validation or AIC for selecting df could be potentially misleading (particularly with high concurvity) since they are designed to choose the df that will lead to optimal prediction of the mortality series” rather than accurate estimation of the coefficient of interest. They also quantified where the bias lives: “For natural splines, the bias drops rapidly between df = 1 and df = 4 per year and is stable afterwards. For penalized splines, the bias drops much more slowly and does not appear to level off until df = 9 or df = 10 per year.” A field whose estimates set national air-quality standards concluded two decades ago that you must not tune the nuisance basis on fit. Media measurement tunes it on fit as a matter of routine.

Three curves on one axis: bias, identifying variation, and prediction error

Same sweep, three things measured. On the left axis, the absolute bias in the search channel's contribution (% of truth) and the identifying variation that survives, meaning the share of the channel's own variance left after the baseline basis is projected out. On the right axis, 5-fold interleaved cross-validated RMSE, the criterion an analyst would actually tune on. Drag the search channel's budget cadence: the slower the budget moves, the sooner the baseline swallows it.

8
Best CV at
|Bias| there
Identifying variation there
Least bias at
CV penalty for it

At the defaults, cross-validated error is minimized at a changepoint every 4 weeks, where the bias is still +22% and only 8.2% of the channel's variance is doing the identifying. The least-biased specification sits at every 8 weeks and pays a 6.8% worse CV score for it, a gap no reviewer would defend a modeling choice on. Note the flat stretch on the right axis from every 16 weeks to every 6: CV improves by 5.8% while the estimate travels nine points of KPI. The criterion has no resolving power exactly where the answer is being decided.

That flat stretch is the honest generalization, and it obliges me to correct a sibling post. Stop Validating Your MMM With Holdout Error argues the general case: a holdout score is a function of the fitted sum, never sees the split, and so cannot rank specifications for a causal estimand. But its limits section conceded that predictive accuracy is “the right criterion for the parts of the model that are forecasts,” listing “baseline flexibility, trend and seasonality structure” among them. That concession is wrong in precisely the case that matters. When time confounds media, the baseline stops being a forecasting component with an incidental causal side effect and becomes the confounder-adjustment surface itself. Prediction error is the wrong criterion for it, for the same reason it is the wrong criterion for the ROI. The concession holds only where spend is uncorrelated with the calendar, which is to say for no real media plan.

None of which licenses the opposite reflex. “Use more knots” is not the lesson, for two reasons the sweep makes visible. The variance cost is real and unbounded: the median interval width grew monotonically across the whole axis, and at the saturated end (one parameter per period in a national model) the media effect is not identified at all. And what minimizes mean squared error is an interior point whose location depends on things you do not know. Over 150 replications, the root-mean-square relative error of the search estimate ran 190%, 150%, 80%, 42%, 37%, 40%, 40%, 46% across changepoint spacings running from 156 weeks down to 2 weeks. That is a genuine interior optimum, at roughly one changepoint every six weeks. Move the budget cadence to quarterly and the optimum moves: the confounder's strength and the channel's spectrum both set it, and neither is known at fit time. There is no defensible \( k \), only a defensible way to handle not knowing it.

Causal ML Named This in 2018

The phenomenon has a name outside marketing, a literature, and even a constructive fix. Hahn, Carvalho, Puelz and He (2018) called it regularization-induced confounding: applying a shrinkage prior to a block of control coefficients biases the treatment-effect estimate, because over-shrinking the controls re-opens the confounding they were included to close. Their answer keeps the shrinkage prior and rewrites the parameterization, so the prior on the nuisance block cannot pull on the treatment effect. The sequel, Hahn, Murray and Carvalho (2020), builds the idea into Bayesian causal forests by splitting the response surface into a prognostic function and a treatment-effect function, so shrinking the first does not shrink the second.

Econometrics reached the same answer from the frequentist side: double/debiased machine learning (Chernozhukov and colleagues, 2018) exists because plugging a regularized nuisance estimate into a naive score contaminates the target with the nuisance's regularization bias. Its Neyman-orthogonal score and cross-fitting are an admission that how you estimate the thing you do not care about determines the bias in the thing you do.

Be precise about what these fixes buy. They remove the bias from regularizing a nuisance function that is otherwise adequate. They do not remove the bias from a nuisance basis that cannot represent the confounder at all. If demand moves on an eleven-week rhythm and your baseline resolves fourteen-week waves, no orthogonalization recovers the difference. The information is not in the model. Both failures are live in an MMM's baseline, and only the first has a technical fix.

⚠️ The orthogonalization that makes it worse

The tidiest-looking fix is to restrict the baseline to the part of the design space orthogonal to media. Let the flexible term explain whatever media cannot, and the media coefficient stops moving with the knot count. Spatial statistics tried exactly this, under the name restricted spatial regression, and has spent a decade retracting it: Khan and Calder (2022) show that inference on the regression coefficient is frequently worse under it than under a model with no spatial term at all. The reason is plain in the media case. When time genuinely confounds media, projecting the baseline onto media's orthogonal complement assigns every scrap of shared time-and-spend variation to media by construction. That pins the estimate to the most credulous answer available and reports a narrow interval around it. Shrink the baseline and you bias toward the confounded answer. Orthogonalize it and you bias toward crediting media with everything the calendar explains.

Meanwhile the field is moving toward more time flexibility, not less. Ng, Wang and Dai (2021) describe the time-varying-coefficient MMM deployed at Uber, in which the media coefficients themselves are weighted over a set of local latent variables so effectiveness can drift. That is a defensible response to real non-stationarity, and it sharpens the problem rather than sidestepping it: a time-varying media coefficient and a flexible baseline are two flexible functions of time competing for the same variation, and their priors decide the split. Every argument above then applies twice.

What the Code Gets Right, and the Gap

This framework has the rule half-implemented, and the missing half is the half this post is about. For named control columns the causal logic is wired in: controls carry a causal role, and the coefficient prior widths are keyed to it. A control tagged CONFOUNDER gets a prior standard deviation of 2.0 on standardized data while an ordinary precision control gets 0.5, and the source spells out why: a confounder “must not be shrunk toward zero” because shrinking it biases the media coefficient by exactly the \( (1-s)\gamma\,\mathrm{Cov}/\mathrm{Var} \) term above. Variable-selection priors are refused on confounders for the same reason, and a deliberately narrow prior on a confounder-tagged control raises a warning that it “can re-open back-door bias on the media effects.”

The time basis sits outside that registry entirely. Its flexibility and prior widths live in TrendConfig and SeasonalityConfig as ordinary hyperparameters (n_changepoints, n_knots, changepoint_prior_scale, spline_prior_sigma, gp_lengthscale_prior_mu), with no causal role attached and no warning when they are tightened. Nothing in the model knows the trend is doing confounder-adjustment work, so nothing objects when it is set to whatever the last fit liked. That is not hypothetical, and the code records the incident:

💡 From the source, on a default that was changed

“The old, much tighter defaults (0.1 / 0.05) effectively pinned the trend near zero and pushed real trend/structural breaks into media and intercept.” (model/trend_config.py). The two numbers are the prior standard deviation on the linear growth rate and the Laplace scale on each Prophet slope change, and both were widened by a factor of five to ten. The seasonality config carries the same warning on its own axis: the Fourier amplitude prior defaults to 0.3 on standardized \( y \), and raising it is advised for strongly seasonal categories, “or the seasonal signal gets squeezed into trend/media.” Two components, two documented instances, one mechanism.

One genuine guard exists, and it is worth distinguishing from what it does not do. In the extension models every non-linear trend family is zero-centered so its level is absorbed by the intercept, and the source calls this “the identifiability guard.” Useful, and unrelated to this failure. A guard on the level does nothing about the shape, and it is the shape (which frequencies the baseline may fit) that decides how much confounding reaches the media coefficients.

Put the Baseline in the Pre-Registered Set

If no setting is defensible and no criterion selects one, then choosing one and reporting its interval is the garden of forking paths with extra steps, a researcher degree of freedom that moves the headline number by a factor of three. The honest move is the one the framework already implements for the transform families: declare a set of defensible specifications before the fit, run all of them, report how the estimand moves. The change argued for here is that the set must include the baseline.

A spec curve on the nuisance

Seven defensible baseline specifications on the same planted world (no trend, linear, three changepoint counts, and two smooth harmonic bases), rendered the way validation/spec_curve.py renders the adstock × saturation grid. One marker per specification within each channel's row, the equal-weight average as a bold diamond whose bars span the set (not a credible interval), and a green tick at the planted truth. Resample to confirm the pattern is not one lucky draw.

Search: spread
Search: truth
Brand: spread
Brand: truth
Equal-weight average, search

The demand-chasing channel spans 11.8% to 37.3% of the KPI across seven defensible baselines, a 3.2× range around a truth of 13.2%. The flighted brand channel spans 6.1% to 6.8% around a truth of 5.3%: nearly invariant to the same axis, and biased about a point high in every one of them. Which channels are baseline-fragile is itself the finding, and it is knowable before any experiment runs.

The per-channel contrast is the useful output. Baseline fragility varies from channel to channel inside a single model, and it maps onto the concurvity number from earlier. The brand channel's three-week bursts live on the fast side of every baseline in the set, so its estimate barely moves. That does not make it correct: it sits about a point above truth in all seven fits, a level error the trend axis cannot fix. The search channel's eight-week budget steps sit on the boundary, so it moves by a factor of three. A six-channel model typically has some of each, and a spec curve says which are which. That is actionable twice over. Fragile channels are the ones whose numbers should never be quoted without an interval spanning the set, and the ones where an experiment is worth its cost, because designed variation is the only thing that moves a channel to the fast side of every baseline you might choose.

Here the machinery exists and the default grid does not use it. SpecVariant already carries a first-class trend axis alongside adstock, saturation, controls and kpi_level, plus a deep-merging overrides escape hatch reaching any spec path including the trend priors. But default_spec_variants() sweeps adstock forms crossed with saturation forms and nothing else, the two axes already known to carry a well-studied identification problem, while the axis that silently rescales every ROI in the report keeps its default in every variant. Declaring it takes five lines.

# illustrative
from mmm_framework.validation.spec_curve import SpecSet, SpecVariant, run_spec_curve

baseline_axis = SpecSet(
    rationale="Baseline flexibility is confounder adjustment, so it is swept, not tuned.",
    registered_at="2026-07-24T09:00:00Z",
    variants=[
        SpecVariant(name="linear",             trend="linear", primary=True),
        SpecVariant(name="piecewise 10cp",     trend="piecewise"),
        SpecVariant(name="piecewise 26cp",     trend="piecewise",
                    overrides={"trend": {"n_changepoints": 26}}),
        SpecVariant(name="spline 20k",         trend="spline",
                    overrides={"trend": {"n_knots": 20}}),
        SpecVariant(name="gaussian process",   trend="gaussian_process"),
        # the prior width is the same knob as the knot count -- sweep it too
        SpecVariant(name="piecewise 26cp tight", trend="piecewise",
                    overrides={"trend": {"n_changepoints": 26},
                               "priors": {"trend": {"changepoint_prior_scale": 0.05}}}),
    ],
)
result = run_spec_curve(base_spec, dataset_path, variants=baseline_axis)

Three details carry the argument. The set is a SpecSet with a rationale and a registered_at stamp, meant to be serialized into the pre-fit design readout before anyone sees a result. The spread is only evidence if the set was fixed in advance. The variants sweep both spellings of flexibility, knot counts and prior width, because the deep dive shows they are one knob. And the run is deliberately not weighted by predictive skill: run_spec_curve defaults to equal weights for the reason its own docstring gives, that two specs can predict the KPI equally well while splitting it very differently between media and baseline. Every claim in this post is a claim about that split.

What This Does Not Establish

It does not establish a number of knots, and any reader who leaves with one has read it backwards. The mean-squared-error optimum in the simulation was interior and moved with the confounder's strength and the channel's budget cadence, neither of which is observable at fit time. The deliverable is a swept, pre-registered, reported axis. Prescribing a value of \( k \) is beyond what this simulation can support.

The demonstration is a demonstration: least-squares fits on a stylized two-channel national world, one latent confounder, Gaussian noise, hinge-and-harmonic bases standing in for Prophet and HSGP, hard degrees of freedom standing in for shrinkage. A Bayesian fit with informative priors will damp some swings and replace the sharp \( k \) with the softer \( s_k \) spectrum. The failure exists, and it is large under conditions that are not adversarial. How often a given team's model lands on the wrong side of it is not something a simulation can say.

I did not demonstrate the mirror-image failure. It is natural to expect a very flexible baseline to absorb the media's own variation and attenuate the estimate toward zero, and I looked: across budget cadences of 8, 26 and 52 weeks, mean bias fell monotonically with flexibility and never went materially negative. What rose monotonically was the median interval width. So this generator supports the bias-variance framing Meridian states, and the hard limit at the saturated end (a national model with one parameter per period identifies nothing) is arithmetic rather than a simulation result. Attenuation may well dominate with heavier adstock, slower channels or shorter panels. I have not shown it.

Geo panels change the calculus in a way the simulation cannot capture, and this is the most important limit. A national model has one row per period, so the time effect and the media block compete for a single series. A geo panel with a period effect shared across regions still has within-week cross-geo contrasts, which is why Meridian can default geo models to a knot per period and why a geo MMM tolerates far more baseline flexibility. Everything above applies most sharply to national weekly models, which is what most brands run.

Finally, the spec curve diagnoses and does not repair. The equal-weight average across the seven baselines lands at 24.0% of the KPI against a truth of 13.2%, because most members of the set are biased in the same direction and averaging biased estimates does not produce an unbiased one. A spec curve is a fragility report. When it comes back wide, the answer is designed variation, an experiment, or an honest interval spanning the set, not a better average. And none of this addresses a confounder that is not a function of calendar time at all (a regional demand shock, a competitor's launch in two markets), where no flexibility setting helps and the argument moves to measured proxies and their own failure modes.

Takeaways

  • Whenever budgets follow the calendar, an MMM's trend/seasonality basis stops working as a forecasting component and becomes a confounder-adjustment surface. Meridian's own documentation transmits the trade-off explicitly: the bias-variance trade-off in time effects becomes one in the causal effects of media.
  • A flexible baseline is algebraically a saturated block of shrunk regressors (Hodges–Reich). Shrinking a confounder by 90% leaves 90% of the bias, \( \hat\beta \to \beta + (1-s)\gamma\,\mathrm{Cov}(X,C)/\mathrm{Var}(X) \), and the knot count and the prior width are two spellings of one knob.
  • The error signal points the wrong way: on a planted world, the stiffest baseline put the demand-chasing channel at 35.0% of the KPI against a truth of 13.2% with the narrowest interval in the sweep, covering the truth in 0 of 200 replications. Loosening it walked the estimate onto the truth and widened the median interval from 6.2 to 15.3 points. Recovery-coverage diagnostics cannot see this, because they simulate from the model.
  • Prediction cannot select the setting. Cross-validated error was minimized where the bias was still 22% and only 8.2% of the channel's variance identified anything, and it was flat over the region where the estimate moved nine points of KPI. Peng, Dominici and Louis established the same for air-pollution models in 2006.
  • The phenomenon has a name, regularization-induced confounding (Hahn et al., 2018), and its fix removes regularization bias, not the bias from a basis that cannot represent the confounder. The tempting orthogonalization repair is restricted spatial regression, known to make inference worse.
  • Ship it as a swept axis, not a tuned one. SpecVariant already has a trend field and overrides reaches the trend priors, while default_spec_variants() sweeps only adstock × saturation. Add the baseline, pre-register the set, report per-channel fragility, and publish each channel's \( 1-R^2 \) against the basis. When that reads 8%, the ROI is an extrapolation however narrow the interval looks.

References