Budget allocation under power laws: from Chinchilla to media mix
Summary
The optimisation transfers completely: all three problems are solved by equal marginal return per unit of the constrained resource, and in each the fitted curvature exponents (scaling exponents, elasticities, Hill slopes, the of sampling) fix the split. The estimation lesson transfers as a warning: Kaplan and Chinchilla fitted the same functional family and disagreed on the allocation ( versus ) because of a biased measurement protocol and a narrow fitting range — exactly the two ways an MMM response curve goes wrong. What does not transfer is the data situation: scaling-law losses are directly observed, nearly noiseless, experimentally designed and stable over six to eight orders of magnitude, whereas media response is a causal counterfactual estimated from about a hundred observational weeks, with carryover, competitor reaction, drifting elasticities and cross-channel interactions that have no counterpart in .
Answer
1. The shared first-order condition
| Chinchilla | Media mix | Single budget level | Experiment design | |
|---|---|---|---|---|
| Objective | incl. post-period to | CI half-width (or EIG) | ||
| Constraint | (a product) | (a sum) | none — level is free | geos, weeks, spend differential |
| FOC | for all | marginal precision per dollar equal across levers | ||
| Exponents give | , | spend shares | ||
| Source | Compute-Optimal Training (Chinchilla) | ROAS, mROAS, and Optimal Media Mix | Optimal Marketing Decisions and Forecasting | Power Analysis and Sample Size, Geo-Experiment Design and Power Analysis |
The Chinchilla note already says the media problem is “solved by the same Lagrange condition of equalized marginal returns per unit of cost”. Two refinements make the correspondence exact.
Synthesis — sum versus product constraints. Chinchilla’s budget is multiplicative, so it is additive in and the condition equalises derivatives with respect to log inputs: , which is the note’s “elasticity-weighted reducible losses … are balanced”. A media budget is additive in dollars, so the condition equalises derivatives with respect to dollars — mROAS. For a constant-elasticity response (^def-power), , i.e.
Budget share is proportional to elasticity times contribution. This is the multi-channel form of Dorfman–Steiner (advertising-to-sales ratio ; ), and of the vault’s multiplicative-mix result that spend ratios across instruments equal elasticity ratios. It also explains why ranking channels on average ROAS misallocates: at the optimum a low-elasticity channel must show a higher ROAS.
Synthesis — the Hill tail is a Chinchilla term. The Hill transform is written in ^betahill-eq as . For the shortfall from saturation is — a power-law deficit with "" and exponent , sitting below an asymptote just as sits above the irreducible . The same parameter trade-off that makes “essentially unidentifiable” applies to fitted far from the asymptote.
How the exponents move the split as the budget grows. Kaplan’s 10× compute buys 5.5× parameters and 1.8× tokens; Hoffmann’s buys about 3.2× of each. In media the expansion path is set the same way: with and power curves, , so less-curved channels absorb a growing share of incremental budget (synthesis). Under Hill saturation every channel’s mROAS eventually falls below cost, which is where the Dorfman–Steiner level condition takes over — a question Chinchilla never asks because is exogenous.
Experiments are the same problem with exponent . says precision is a power law in sample size. The geo notes give the multi-input version: TBR’s iROAS half-width falls as in spend intensity, as in pretest length but is “bounded below” by , improves with test length “with diminishing returns”, and grows with cooldown length at fixed spend. Synthesis: that floor is the Kaplan phrase “when not bottlenecked by the other two” in another domain — pretest weeks cannot substitute for test weeks, just as parameters cannot substitute for tokens. The Chinchilla and scaling-law notes draw the same link: IsoFLOP profiles are “the analogue of a power curve traced by pilot simulation”, and the pseudo-geo-experiment procedure is that pilot simulation.
2. The estimation problem: Kaplan versus Chinchilla as a cautionary tale
Both papers assumed power laws. They disagreed on the allocation because of how the curve was measured.
| Failure in scaling laws | What the vault records | MMM counterpart |
|---|---|---|
| Protocol bias | Kaplan used one cosine schedule for all horizons; intermediate losses overestimate what a matched schedule achieves, which “understates the value of training smaller models on less data” | Mis-specified carryover: response read before adstock has played out understates a channel’s return (stated in the Chinchilla note). ROAS must sum to |
| Narrow range | Most Kaplan runs below 100M parameters; “slight curvature in the FLOP-loss frontier”, so “extrapolating from small models misleads” | outside the observed spend range is unidentifiable; the model “cannot extrapolate”. Budget 0.5 (sparse region) gave a three-mode posterior for the optimal mix; budget 1 gave a tight unimodal one |
| Exponent leverage | Rounded published constants give B at Gopher’s budget, the unrounded fit 40B; “the third decimal of the exponent moves the answer materially”. Kaplan’s crossover point is uncertain by an order of magnitude | Two-year samples bias Hill at by to ; posteriors of and in the shampoo data roughly equal their priors |
| Fit dominated by the dense regime | Huber loss needed; larger “pushes the model to overfit the small compute regime and poorly predict held-out data from larger runs” | Synthesis: most weeks sit near average spend, so the likelihood is dominated by exactly the region that says least about curvature |
| Bookkeeping | Embedding parameters counted or not changes the trend | Spend versus impressions versus GRPs; change-period definition |
What Hoffmann et al. did about it is the transferable method: three independent estimation routes (training-curve envelope, IsoFLOP profiles, parametric fit) agreeing on ; bootstrap intervals on the exponents; a head-to-head test at FLOPs; and one confirmatory run at scale (Chinchilla 70B beats Gopher 280B, 67.6% versus 60.0% MMLU). Note that Approach 2 involves no extrapolating functional form: hold the budget fixed, vary the split, find the valley.
Synthesis — the iso-budget geo test. The media analogue of an IsoFLOP profile is a multi-cell geo experiment at fixed total spend with different channel splits across cells, fitted with a parabola in log-share. It measures the allocation optimum directly, inside the range where it will be used, and is the natural triangulation partner for an MMM-derived mix.
3. Deciding under an uncertain curve
- Propagate, do not plug in. The ROAS note is explicit that averaging parameters and then plugging in is wrong; push each posterior draw through the metric. The scaling-law note recommends the same upgrade: a Bayesian fit with posterior uncertainty “carried through to the design decision”.
- Route A is the decision; route B is the diagnostic. Decision Analysis gives . Optimising the average predicted sales over draws (Eq. 15) is that rule and gives one stable mix. Optimising per draw (Eq. 16) gives the posterior of the argmax, whose multimodality tells you whether the data determine the answer at all.
- Flat optima cut both ways. Kaplan: models between 0.6× and 2.2× the optimal size cost only 20% more compute. Jin et al.: variance in estimated sales across draws is comparable to or larger than the variation caused by changing the mix, so the optimum “is not trustworthy”. Synthesis: these are the same fact — a flat objective makes the argmax ill-determined but deviations cheap. Report expected regret of the chosen mix against each draw’s optimum, not only the argmax distribution.
- If the argmax posterior is diffuse, buy information. In Bayesian Optimisation the multi-modal optimal-mix posterior is and the location-information loss targets its entropy; Expected Information Gain is the design objective when the deliverable is the curve itself. Both have diminishing returns, so the experiment budget is a fourth allocation problem nested inside the third. The BO note’s defence of myopia — “the surrogate is often wrong” — applies with more force to an MMM than to a GP.
4. Where the analogy fails
- The objective is not observed. Pretraining loss is measured directly with seed noise of about 0.02; incremental sales are a counterfactual. BIC “select[s] the best regression model, not the best causal model”, and the scaling-law runs were designed experiments while MMM spend is observational.
- Dynamic range. Eight orders of magnitude in compute versus whatever spend variation happened to occur in weekly observations. Lengthening the history is “not the recommended fix” because market conditions drift; the vault’s fix is pooling brands or geos for informative priors.
- Carryover and timing. Compute has no adstock. Media has retention , delayed peaks , a long-run effect about double the short-run, a 90% duration of six to nine months, and possible hysteresis. The allocation is over a flight pattern: concave response implies even spending, S-shaped response can make pulsing optimal, convex response gives corner solutions (Shape of the Marketing Response Function). With the objective is not concave and the Lagrangian condition is necessary, not sufficient.
- Competition. Own-response is “a partial equilibrium result”; rivals react and “ignoring reactions overstates the value of marketing investments”. Nothing reacts to your FLOPs.
- Non-stationarity. Scaling exponents transfer across text distributions with a constant offset. Advertising elasticities decline over the life cycle (0.625 → 0.496 → 0.274; new products 0.26 versus established 0.05), so last year’s curve is a prior, not a law.
- Interactions. Jin et al.’s ROAS machinery assumes additive media effects with no cross-channel spillover; the multiplicative form builds synergy in; price and non-price advertising move price sensitivity in opposite directions. is additively separable by assumption, and interaction enters only through the constraint.
- Costs outside the objective. Inference cost scales with , which pushed Chinchilla to a smaller model than loss alone required; margin, long-run brand effects and CLV play that role in media.
- One shot versus repeated play. “It is typically only feasible to train these large models once”, so extrapolation is unavoidable. Media is re-allocated weekly, so sequential learning can replace extrapolation.
Practical Implications
- Allocate on mROAS, never ROAS, computed per posterior draw including the carryover tail. Sanity-check with against the meta-analytic –.
- Do not recommend a mix outside the observed spend range without an experiment. Flag any channel whose posterior sits at the edge of its range-constrained prior.
- Audit the measurement protocol before the curve — adstock window , change-period definition, cooldown length. That, not the functional form, is what separated Kaplan from Chinchilla. Residual autocorrelation to lag 15 in the shampoo model is the kind of signal to chase.
- Triangulate: parametric MMM, an iso-budget multi-cell geo test, and a confirmatory holdout of the recommended mix.
- Report three numbers: the route-A mix, the route-B spread, and expected regret. If regret is small, stop optimising; if the spread is wide and regret is large, spend on information.
- Size experiments with the same logic: spend intensity first (half-width , watching saturation), then test length, and only then pretest length, which hits a floor; do not pad cooldown.
- Prefer parsimonious curvature (reach, ; geometric adstock) when likelihoods tie, as BIC did, and carry functional-form uncertainty as a multiverse rather than as extra parameters.
- Treat fitted curves as perishable and re-estimate on a schedule.
Source Notes
| Note | Relevance |
|---|---|
| Compute-Optimal Training (Chinchilla) | Constrained problem, closed-form frontier, three estimation approaches, why Kaplan differed, the marketing paragraph |
| Neural Scaling Laws | Power-law premise, Kaplan allocation, flat optimum, where the laws must fail |
| ROAS, mROAS, and Optimal Media Mix | Metrics, budget-constrained problem, routes A/B, three-mode posterior |
| Shape (Saturation) Effects, Bayesian Estimation and Priors for MMM, MMM Model Selection and Application | Hill form, unidentifiability, no extrapolation, small-sample bias, BIC, bimodal optimum |
| Optimal Marketing Decisions and Forecasting, Functional Forms in Marketing, Shape of the Marketing Response Function | Dorfman–Steiner, elasticities, shape-dependent pulsing |
| Advertising and Promotion Effects | Elasticity benchmarks, long-run multiplier, life-cycle drift |
| Carryover (Adstock) Functional Forms, Reaction Functions and Competitive Dynamics | Where the analogy breaks |
| Power Analysis and Sample Size, Geo-Experiment Design and Power Analysis, TBR Design Sensitivity and the Stationarity Assumption | Design as allocation; scaling of CI half-width in each lever |
| Bayesian Optimisation, Expected Information Gain, Decision Analysis | Acting and experimenting under an uncertain curve |
Related Concepts
- Carryover Effects and Distributed Lags — Koyck dynamics behind the long-run multiplier
- Multivariate Persistence and Cointegration — hysteresis and permanent effects
- Acquisition Functions and Value Loss and Entropy Search — concrete rules for choosing the next spend level
- Type S and Type M Errors — what under-powered experiments do to the elasticities fed into the allocation
- Prior Predictive Checking — check implied response at target spend before trusting an extrapolation
- Q - BED vs Bayesian Optimization vs Bandits for Media Experimentation — which objective to use once you decide to buy information
Gaps
- No note derives the multi-channel allocation rule; the share formula and the Hill-tail correspondence are synthesis from the vault’s definitions.
- Source-note correction (resolved 2026-09-18): while answering this question, Optimal Marketing Decisions and Forecasting was found to state with for the multiplicative model, and its Dorfman–Steiner derivation lines were garbled. The first-order condition yields ; the source note has been corrected accordingly. Still worth checking against Hanssens et al. Ch. 9 for the book’s own notation.
- No value-of-information note: nothing on EVPI/EVSI or on how much media budget to divert to experiments, and Decision Analysis is a short stub.
- No robust or risk-averse allocation under posterior uncertainty, and no dynamic allocation with adstock state beyond the optimal-control sketch.
- Post-Chinchilla scaling work is absent (inference-aware or data-constrained multi-epoch laws), so point 7 of Section 4 rests on one sentence of the Chinchilla note.
- Competitive-response-adjusted response curves are not connected to the Bayesian MMM notes.
Follow-Up Questions
- How many cells and weeks does an iso-budget geo test need to locate the optimal split to within a given regret?
- What is the expected regret of the route-A mix in the Jin et al. Scenario II simulation, and does it justify the label “not trustworthy”?
- How does the optimal allocation change when the state is adstock and the objective is discounted CLV?
- Can a hierarchical prior over Hill exponents across brands play the role that cross-scale data played for Chinchilla?