Your Curve Assumed You Wouldn't Act On It
Every fitted marketing mix model ships an optimizer that treats its response curve like a fixed law: plug in a budget, differentiate, report the argmax. A response curve summarizes what happened under one particular spending policy, the one your team actually ran. No model is ever a finished, perfectly correct picture of reality. It is a simplification built to be useful for one decision at a time, and useful only as long as the world underneath it still looks roughly like the world it was fit on. Change the policy, and you are no longer reading the curve. You are asking it a question it was never fit to answer.
This framework already has one instrument built for a version of that problem: a per-channel max_obs_multiplier check that flags a recommendation pushing a channel past anything its own history ever showed. That check is real and well built. It is also not the subject of this post. A recommendation can clear every channel's own historical range, trigger zero flags, and still lean on a spending pattern the data never produced, because the curve was fit under one joint policy running across channels and time together rather than under independent per-channel bands. In the simulation below, two channels that always moved together produce a reallocation that passes both individual range checks while landing 7.6 standard deviations outside the joint history the curve was ever estimated on. That is the sharper question worth learning to ask of the model, and this post is an invitation to start asking it.
A Fact About a Policy, Not a Law About the World
A marketing mix model's response curve is estimated from whatever spending actually happened, and that spending was never exploratory. A media plan is a negotiated, calendar-locked commitment, so the typical channel runs at close to the same weekly level for most of the year, with an occasional flight or promotional bump on top. This framework's own defaults quietly assume as much: a client report that hasn't been through the Planner studio caps its "default reallocation" at ±20% of current spend per channel, and a plain-language comment beside the per-channel headroom check spells out the edge case. A channel that has always run at a perfectly constant level has nowhere to scale up without counting as extrapolation.
The curve fit on that narrow, steady history is then handed to an optimizer that treats it as portable. Budget in, argmax out. Nothing in that step asks whether the curve survives the move it's about to recommend. Economist Robert Lucas asked the same question of macroeconomic models back in 1976, in a paper that became one of the field's most durable warnings. His point, in plain terms: a relationship estimated from historical data bakes in the decisions people were making under the old rules, so you can't trust it once the rules change, because the relationship itself was a product of those rules. Lucas was writing about consumption functions and Phillips curves. A media response curve is the same kind of object wearing different variable names, a pattern that held while certain behaviors stayed fixed, now being asked to keep holding once one of those behaviors changes on purpose.
Marketing researchers Van Heerde, Dekimpe and Putsis carried the critique into marketing econometrics directly. Their 2005 paper lays out exactly when a marketing model's parameters are themselves tied to the policy that generated the data, and it notes that the point had gone largely unnoticed in marketing despite applying just as strongly there as in macroeconomics. A budget optimizer that maximizes a fitted curve and calls the argmax "the recommendation" is, two decades later, doing the thing they were flagging, with a Bayesian saturation curve standing in for a consumption function.
Definition: a policy-conditional response curve
In plain terms: the curve you fit describes how the channel behaved under one specific spending pattern, and swapping in a new pattern is exactly what an optimizer's recommendation asks you to do. Written more formally, a fitted response curve is really \( \hat f_{\pi}(\cdot) \), not \( \hat f(\cdot) \): a function estimated from data generated under one joint spending policy \( \pi \) across every channel and week. An optimizer that solves
$$ \mathbf{s}^{\star} \;=\; \arg\max_{\mathbf{s}} \sum_{c} \hat f_{\pi_0}(s_c) \quad \text{subject to} \quad \sum_c s_c = B $$and reports \( \mathbf{s}^{\star} \) as the new plan is quietly assuming \( \hat f_{\pi_0} = \hat f_{\pi_1} \), where \( \pi_1 \) is the policy that acting on \( \mathbf{s}^{\star} \) creates. A standard fit never tests that assumption. And \( \mathbf{s}^{\star} \) sitting inside each channel's own historical spend range doesn't confirm it either: \( \pi \) is a statement about the JOINT pattern across channels and time together, and a point can sit comfortably inside every individual channel's range while sitting far outside the combination those channels ever formed. A person can be an average height and an average weight without having an average build.
Performative Prediction, and the Older Critique It Formalizes
The Lucas critique is fifty years old and speaks in the language of macroeconomics. A cleaner, more current version of the same idea comes from machine learning research. Computer scientists Perdomo, Zrnic, Mendler-Dünner and Hardt gave it a name in 2020: performative prediction. A prediction is "performative" when acting on it changes the very thing the prediction was trying to describe. Their intuitive fix is to retrain until the model and the world it produces agree, but that only counts as stable if the agreement actually survives being deployed again. A budget optimizer that maximizes a fitted curve and hands you the argmax is doing exactly this: the recommendation is the action, and the curve was fit on data generated before that action existed.
Their paper is about loan-scoring classifiers and content recommenders, not media spend, and the reason to bring it up here is narrower than "marketing invented this problem." A fifty-year-old critique of economic policy models and a five-year-old idea from machine learning describe the same underlying pattern, independently, in fields that otherwise share almost no methodology. The convergence is structural. Any model asked to justify a decision that changes the world the model was fit on runs into some version of it, and a maturing measurement practice keeps rediscovering that lesson, each time as a chance to build a sharper check.
One distinction sits underneath all of this. Acting on a prediction can move the world past what the model has seen in two different ways. The first is the easy, catchable one: the action pushes some input past anything the training data ever showed on that input's own axis. That is the univariate case, precisely what this framework's max_obs_multiplier is built to flag. The second is subtler, and it's the one the next two figures explore. The action can leave every individual input exactly where the training data always put it, while moving the combination of inputs into territory the training policy never visited together. A regime can change this second way without tripping a check designed to watch one input at a time. That check is doing what it was named to do, and the next layer of the work simply belongs somewhere else.
The Curve That Looked Perfectly Calibrated
Start with the single-channel version, because it is the version this framework's own extrapolation check is built to catch, and on its own terms does catch. Picture a channel run at a near-constant weekly level for two years, drifting ±8% around its mean on a rhythm set by the planning calendar. Now suppose sales were also moved by something that isn't media at all, a competitive push or a category tailwind, and that this something happened to track the channel's own week-to-week wobble for no causal reason whatsoever. Both were simply driven by the same seasonal planning cycle that set the channel's spend in the first place. Fit the standard single-channel regression on two years of that history and the coefficient absorbs both effects, because within the historical range the confound and the channel's own spend are indistinguishable to a regression that has never seen them move apart. At the figure's default comovement strength, the fitted coefficient comes out 33% above the channel's true, causal per-dollar effect. The fit does not confess it. Across the full two years in-sample, the residual is almost invisible, because the confound rode along with every historical dollar the model ever saw.
A channel run at a near-constant level, confounded by something that used to move with it
Left: two simulated years of weekly sales attributable to one channel, observed (including the confound) against what a standard single-channel regression fits. Right: the same channel's response curve extended by spend multiplier, plotting the fitted model's prediction against the true, causal response once the confound is decoupled from a deliberate spend increase. The dotted vertical line is the channel's own max_obs_multiplier ceiling. The diamonds mark the optimizer's actual recommendation, comfortably inside it.
At the default comovement strength, the fitted coefficient (132.8) overstates the channel's true causal effect (100.0) by 33%, and the in-sample fit over two full years of history is close to exact, because the confound rode along with every historical dollar the model saw. The optimizer's recommended multiplier, 1.06× current spend, sits comfortably inside the channel's own max_obs_multiplier ceiling of 1.08×, so no extrapolation flag fires, and the predicted incremental return there still overstates the true, decoupled response by 33%. Drag comovement to zero and the two curves coincide exactly. Drag it to its maximum and the overstatement reaches 55%.
Now push spend to the level the optimizer actually recommends off that curve. It sits inside the channel's own headroom, comfortably below its max_obs_multiplier, with no flag raised anywhere in the report. But the confound doesn't come along: it was never caused by this channel's spend, only correlated with the planning calendar that set spend too, and scaling this channel up on purpose breaks that link. Breaking that link is the whole point of a deliberate move, rather than waiting for the calendar to shift things on its own. The true incremental return at that spend level is the causal coefficient alone. The fitted curve, still carrying its inflated slope, overstates it by exactly the bias baked into the historical fit: a third, at the figure's defaults, and worse the harder the two series moved together historically.
In Range on Every Axis, Off the Map
The single-channel story already shows the extrapolation check's blind spot, but it still needed an unmeasured confound to do the damage. The two-channel version needs nothing hidden at all. It's a property of the joint history of two channels the model measures directly, which makes it the sharper illustration of why "inside every channel's own range" and "inside the region the curve was estimated on" are two different claims.
Picture two channels run under one shared media plan for two years, moving together tightly, though never perfectly. The same planner set both budgets against the same calendar, so a high-spend week for channel A was very often a high-spend week for channel B too. max_obs_multiplier checks each channel's own axis: has this channel's recommended average spend stayed below the largest single week it actually ran? Both channels can pass that check independently while the pair of them, taken together, describe a spend combination the historical data never remotely produced. The check simply has no way to see the other axis, and that is exactly the opening for a second, complementary instrument, one that looks at the two channels together instead of one at a time.
Two channels that always moved together
104 simulated weeks of two channels' spend, run under one shared media plan. The dashed box is each channel's own historical min/max, which is all a per-channel check like max_obs_multiplier can see. The diamond is the optimizer's recommended reallocation: more to channel A, funded by pulling back on channel B.
At the default co-movement setting (realized sample correlation 0.86), the recommendation sits at 1.38× channel A's mean spend against a channel-A ceiling of 1.41×, and at 0.76× channel B's mean (a cut, which can never extrapolate upward). Both rows would read within_observed_range: true in this framework's own optimizer output. The same point measured against the two channels' JOINT history is 7.6 standard deviations away, a spend configuration the historical data never remotely produced. Drag co-movement down to 0.5 (sample correlation 0.54) and the distance falls to 4.2 SD, still nowhere near a boundary either per-channel check would ever see.
At the default comovement strength, the reallocation the optimizer recommends (more to channel A, funded by pulling back on channel B) sits inside channel A's own observed ceiling and, since a cut can never extrapolate upward, trivially inside channel B's too. Both rows of the per-channel table this framework prints would read within_observed_range: true. Measured against the joint history instead of either channel on its own, that same point is 7.6 standard deviations from anything the two channels together ever did. Under the historical pattern, that is close to impossible. The two channels moving somewhat independently is exactly the property the historical data doesn't have, and a check that scores each channel's own axis one at a time has no way to register that absence.
Two Failures This Is Not
Two adjacent critiques of budget optimizers are easy to mix up with this one. Telling them apart matters, because the fix for each is different, and chasing the wrong fix spends real effort in the wrong place.
The first comes from this framework's own earlier post on adstock and scheduling, and the two are complementary rather than competing. That post covers the objective side well: even handed a perfectly estimated response curve, is a window-summed total the right thing to maximize for a channel whose effect persists past the window, and is a flat spending plan right when the fitted curvature says otherwise? It's explicit that its results assume "the true response surface with zero uncertainty." The curve itself is never in question there, only what you do with it. This post is the complementary piece: how the world underneath that same curve can shift once you act on it, even when the curve and objective were both exactly right at the time. Push a little further on that post's own caveat that its results ignore a competitor's reaction, and you land close to this post's point: a rival reacting to your move is one way the regime can change. A co-moving confound decoupling, or a joint spending pattern breaking apart, are others. "The curve doesn't know about competitors" turns out to be one instance of a broader truth: the curve doesn't know about anything it assumed would stay fixed.
The second adjacent critique targets the statistics of the argmax rather than the curve itself: optimize over a noisy estimate and the winning plan gets picked partly for being lucky, so its projected uplift runs optimistic and its composition shifts across posterior draws. This framework already takes that seriously. It re-optimizes per posterior draw, reports the probability the plan beats the status quo, and prices the expected KPI left on the table from committing to one plan under parameter uncertainty. All of that accepts the fitted curve as the right target and asks how confidently the argmax was found on it. This post's point sits at a right angle to that one, and it survives even a perfectly known, zero-uncertainty curve: give every posterior draw the identical, noiseless true response function, and every draw still shares the same now-invalid regime assumption, so re-optimizing per draw doesn't touch it. A well-calibrated posterior and an honest decision-uncertainty report are entirely compatible with a plan that's confidently optimal for a curve that no longer applies.
What the Framework Already Catches, and What's Worth Adding Next
Give the optimizer credit for the check it already runs, honestly and by name. This framework's per-channel max_obs_multiplier (in planning/budget.py) computes, for each channel, the largest single-period spend the model actually saw divided by that channel's mean spend. That ratio is the biggest multiple of current spend the historical data ever supports. When a recommended multiplier clears that ceiling, the optimizer flags the channel as extrapolated, widens its reported spend interval for the extra uncertainty of operating past observed spend, and prints a note recommending a test before committing to large scale-ups. That is a real and honestly worded defense against gross, single-channel extrapolation. The default client-facing reallocation goes further still, keeping every channel within ±20% of its own historical band by construction, precisely so the check rarely even has to fire.
The exact mechanics
ResponseCurves.max_obs_multiplier computes "the largest current-spend multiple that stays within observed support." optimize_budget checks every channel's recommended multiplier against it and, when channels clear it, prints that they are "recommended beyond the spend range the model has observed" and that the reader should "confirm with a test before committing large scale-ups." The curve-sampling function's own docstring separately notes that "the model is additive in channels, so each channel's curve is unaffected by the others' scenario spend." That is accurate about how the posterior-predictive pass is computed, and it is a useful pointer toward where a joint-support check belongs.
What this check doesn't yet do is look across channels together. It works per channel and per axis, one at a time, and the response curves it's built on are generated the same additive way, so nothing in the pipeline looks at how channels moved together historically. A recommendation can be entirely within every channel's own historical range (clearing max_obs_multiplier on every row, triggering zero notes, exactly what the figure above showed) and still describe a joint spend combination nothing in the training data ever exhibited. That's the natural next instrument worth building: something that asks whether the recommended point sits inside the convex hull, or within some bounded distance, of the joint historical spend distribution. Nothing like it exists in this codebase yet, and it would be a useful layer to add.
A second instrument, not a fix to the first
The per-channel check does exactly what its name promises. What's described here is a second, complementary question rather than a defect in the first. "Is this channel's own spend in range" and "is this joint spend configuration in range" are genuinely different checks, and having the first doesn't hand you the second for free. A clean within_observed_range reading clears one necessary condition. That is worth having, and it is a good reason to build the second layer on top rather than to call the question answered.
Treat the Curve as a Hypothesis, Not a Fact
None of this is an argument for refusing to act on a fitted curve. A curve fit honestly on real spend variation is still the best evidence a team has, and setting it aside in favor of gut feel is its own kind of error. The better response is smaller: stop treating the argmax as a finding, and start treating it as a hypothesis about the new regime, one sized so a small, controlled move can confirm or refute it before the full recommendation gets trusted.
This framework already has a module built around exactly that posture, worth naming because its whole design leans into the idea rather than working around it. Instead of fitting a response surface once and handing it to an optimizer, continuous_learning/ keeps re-learning the surface after every action, through small designed interventions folded back through a closed loop that carries the posterior forward wave by wave. It is not a complete fix for the failure this post describes, since it was built for geo-experiment programs rather than a national MMM's one-shot reallocation. But its design principle is the right one to borrow: don't ask one fit to survive an action it never saw, ask a sequence of small, deliberately decoupled moves to re-estimate the curve under the regime the action creates. A companion piece answers a related question one level upstream. Would this schedule actually move the channel's structural parameters off their priors? Those are the same adstock and saturation parameters this framework's other posts keep returning to as the hardest things in an MMM to pin down.
Where this lives
The closed-loop machinery is continuous_learning/design.py::central_composite for the designed interventions, folded through continuous_learning/loop.py's LearningState and run_closed_loop to carry the posterior across waves. The upstream identifiability question is answered by planning/identification.py's Fisher/Laplace bound.
Committing once, versus testing in stages
The same confounded single-channel world as the first figure. "One-shot" commits fully to the optimizer's recommendation and never revisits it. "Staged" runs three small, deliberately decoupled test moves (spend levels the historical calendar never produced on its own) and refits after each one.
At the default comovement strength, both paths start 33% high. Committing to the one-shot plan and never testing it leaves the estimate exactly there, indefinitely, because nothing in a single fit tells the model it was wrong. Three staged test waves of eight weeks each, spending at levels the historical plan never visited on its own, pull the estimate to within 4% of the truth by the third wave. That recovery comes from the only kind of data (generated under the new regime rather than the old one) that can tell a policy-conditional curve it no longer applies.
What This Does Not Establish
A few things, stated plainly so the argument isn't over-driven.
This isn't a case against one-shot budget optimization, or against MMM-based allocation in general. A curve fit honestly on real historical variation is evidence, and setting it aside in favor of intuition is not the safer choice. The claim is narrower: an argmax computed once and acted on in full is a hypothesis about a regime that doesn't exist yet, not a finding about the one that does.
The simulations are illustrative constructions, not a measurement of how often real portfolios run into this. The comovement and cross-channel correlation strengths sit on a slider on purpose. How common either pattern is in a given business is an empirical question this post doesn't answer. A category with genuinely independent channel budgets, or a historical calendar that didn't tie media to anything else driving sales, would show a much smaller effect, possibly none.
A joint-support distance is a diagnostic, not a fix. Flagging that a recommendation sits far from the joint historical distribution says the curve's validity there is unverified. It doesn't by itself say which direction the true response actually points, and no such check is built into this codebase today. The two-channel figure makes the case for building one rather than demonstrating one already closing the loop.
Sequential re-testing has its own failure modes worth watching for. A designed test move too small to decouple from the confound teaches the model nothing. One that ignores adstock memory, or invites a competitive response, reintroduces problems this framework's other posts already cover. Three waves converging cleanly in a controlled simulation isn't a claim about how many waves a real program needs, and continuous_learning/'s own design machinery is built for geo experiments specifically, not for retrofitting every national MMM decision.
The Rule for a One-Shot Reallocation
The practical version is short. Before acting on a fitted curve's recommendation, check what the model can already tell you, and hold one more question in mind that it can't yet answer. It can tell you whether every channel individually clears its own historical range: read the within_observed_range flag and the extrapolation note, and take them seriously, because they catch a real failure mode. It cannot yet tell you whether the recommended point is anywhere near the joint spend pattern the data actually produced, because nothing in the pipeline checks that today. Treat "every channel individually in range" as necessary, never as sufficient.
Size the first move accordingly. The larger the distance between the recommendation and the historical joint pattern, the more it should be read as a hypothesis about a regime that doesn't exist yet, and the smaller the first real step should be. Fund a test at a fraction of the recommended change, large enough to break the historical co-movement rather than ride along with it, before committing to the full reallocation. None of that comes from distrusting the model. It is the same discipline the model already applies to its own historical fit, extended to the one thing a single fit can never verify about itself: whether the world it describes is still the world you're about to create.
Takeaways
- A fitted response curve is really \( \hat f_{\pi}(\cdot) \), a function of the JOINT spending policy \( \pi \) that generated the data, not a fixed law about the channel. Acting on its argmax quietly assumes the new policy leaves the curve unchanged, an assumption nothing in a standard fit tests (Lucas, 1976; Van Heerde, Dekimpe & Putsis, 2005; formalized as performative prediction by Perdomo et al., 2020).
- This framework's own
max_obs_multiplierhonestly catches gross, per-channel extrapolation and prints a note when it fires, a real and well-built defense. The natural next layer looks at channels jointly: a recommendation can sit inside every channel's own historical range and still land far outside the joint pattern those channels ever moved in together. - In a two-channel simulation, a reallocation that clears both channels' individual range checks with room to spare sits 7.6 standard deviations from the joint historical spend distribution at the default co-movement strength.
- A single unmeasured confound that historically co-moved with one channel's near-constant spend inflates that channel's fitted coefficient by 33% at default settings (invisibly, since the in-sample fit is essentially exact), and the bias fully survives inside the channel's own observed spend range.
- This complements rather than repeats this framework's earlier post on objectives and scheduling (which covers the case of a correct, static curve throughout), and sits at a different angle from the optimizer's-curse family of decision-uncertainty critiques (those assume a noisy estimate of a fixed, correct curve). It survives an exactly known, zero-uncertainty response function, and per-draw re-optimization doesn't touch it.
- The constructive direction already exists in this codebase, for a different kind of decision:
continuous_learning/re-learns the response surface after small designed interventions rather than fitting once and optimizing forever. In simulation, three staged test waves take the same starting bias from 33% down to under 4%, where committing to the one-shot recommendation leaves it exactly where it started.
None of this is a verdict on the model. It describes what any fitted model has always been: never a finished, perfectly correct account of the channel, only ever a tool built for one decision at a time. Nothing here is broken and waiting for a patch. Adding a joint-support check, and treating every reallocation as a hypothesis worth a small test first, is the next increment in the same ongoing conversation with the model that got it this far.
References
- Lucas, R. E. (1976). Econometric Policy Evaluation: A Critique. Carnegie-Rochester Conference Series on Public Policy, 1, 19–46.
- Van Heerde, H. J., Dekimpe, M. G., & Putsis, W. P. (2005). Marketing Models and the Lucas Critique. Journal of Marketing Research, 42(1), 15–24.
- Franses, P. H. (2005). On the Use of Econometric Models for Policy Simulation in Marketing. Journal of Marketing Research, 42(1), 4–14.
- Perdomo, J. C., Zrnic, T., Mendler-Dünner, C., & Hardt, M. (2020). Performative Prediction. Proceedings of the 37th International Conference on Machine Learning, PMLR 119, 7599–7609.