Measuring the Long Run
Every budget decision is a bet on the long run. You spend today to build a customer who stays, a brand that carries pricing power, a base of demand that compounds for years. But almost every instrument we have — the two-week lift test, the quarter of sales an MMM sees, the immediate conversion an ad platform reports — measures the short run. The uncomfortable gap between what we decide about and what we observe is the central problem of long-horizon measurement, and it has a name in the causal-inference literature: surrogacy. A surrogate is a short-term quantity you observe now, standing in for a long-term outcome you cannot yet see. Used well, surrogates let you steer toward durable value without waiting years. Used badly, they let you optimize confidently in exactly the wrong direction.
The Gap Between Decision and Data
Consider what a marketer actually wants to know: if I raise brand spend by a million dollars this quarter, what is the effect on revenue — not this month, but over the customer lifetimes I am seeding? The honest answer arrives on a timescale that no experiment tolerates. Lifetime value accrues over years; brand equity shifts over multiple purchase cycles; retention cohorts reveal themselves only as they churn or don't. Meanwhile the decision must be made this planning cycle, and the data in hand covers weeks.
The reflexive fix is to optimize the short-term outcome you can see and hope it tracks the long one. Sometimes it does. Often it does not, and the divergence is not random noise but structural. A promotion that pulls forward demand lifts this month's sales while cannibalizing next quarter's. An aggressive ad load raises immediate click revenue while training users to ignore ads — the phenomenon Google researchers named "ads blindness" (Hohnhold, O'Brien, and Tang, 2015). A discount-heavy acquisition channel books cheap conversions that churn before they ever pay back. In each case the short-term proxy points one way and the long-run outcome points the other. Optimizing the proxy is then worse than optimizing nothing: it applies real pressure toward a goal you did not choose.
This is why "just use a proxy" is not a solution but a hypothesis — one that has to be stated formally and checked. The surrogate literature, born in clinical trials where waiting to observe mortality is both slow and ethically fraught, gives us the language to state it and the tools to check it.
What a Valid Surrogate Demands
The foundational treatment is Prentice (1989), which asked exactly the media question in medical dress: when can a comparison of treatments on a short-term marker substitute for the comparison on the true endpoint? His answer is a criterion strict enough that it still disciplines the field. A surrogate is valid, in Prentice's operational sense, when a test of "no treatment effect on the surrogate" is also a valid test of "no treatment effect on the true endpoint" — which requires the surrogate to fully capture the treatment's influence on the outcome.
Definition: The Prentice full-mediation criterion
Let \( W \) be the treatment (say, a spend change), \( S \) the surrogate observed early, and \( Y^{L} \) the true long-run outcome. The core Prentice condition is that the true endpoint is conditionally independent of treatment given the surrogate:
$$ Y^{L} \;\perp\; W \;\mid\; S. $$In words: once you know the surrogate, learning whether the unit was treated tells you nothing more about the long-run outcome. The entire causal effect of \( W \) on \( Y^{L} \) is routed through \( S \) — the surrogate is a complete mediator. Prentice's paper is often summarized as four criteria (treatment affects the surrogate; treatment affects the true endpoint; the surrogate predicts the endpoint given treatment; and full capture), but the last is the binding one and the others are its scaffolding.
Prentice validity demands that all of treatment W’s influence on the long-run outcome Y flow through the surrogate S (the black path). The red-dashed edges are the killers: a direct path W→Y the surrogate never sees, or an unmeasured U confounding the S–Y association. Either one voids the criterion — and is exactly what the surrogate paradox exploits.
The strength of this condition is also its problem. Full mediation says there is no path from treatment to the long-run outcome that bypasses the surrogate. A single short-term metric almost never earns that. Immediate conversions do not capture the brand-memory effects of an impression; this month's sales do not capture the habit a customer is forming. Prentice himself was cautious about how rarely real markers meet the bar, and the subsequent literature has been a long catalogue of ways single surrogates fail. The failure has a sharp form worth its own section.
The Surrogate Paradox
The most alarming failure mode is not that a surrogate is uninformative — it is that a surrogate can be strongly, correctly associated with the outcome and still point the decision the wrong way. This is the surrogate paradox, and it is not a curiosity but a live risk in any optimization loop.
⚠️ A treatment can improve the surrogate, harm the true outcome, and lie the whole time
Formally (VanderWeele, 2013): a treatment may have a positive average effect on the surrogate, the surrogate may have a positive effect on the true endpoint, and yet the treatment may have a negative effect on the true endpoint. Every pairwise association you can measure looks reassuring, and the net effect is the opposite of what those associations imply. Marginal correlation between surrogate and outcome — the thing dashboards actually show — is no protection at all.
The proxy agrees early, then betrays you
Treatment raises a short-term surrogate the dashboard tracks (blue). Early on the true long-run outcome (the other line) rides right along with it, so the surrogate looks like a faithful stand-in. But a slow hidden path — eroding trust, pulled-forward demand, churny acquisitions — accrues over the horizon. Turn it negative and the surrogate keeps climbing while the truth turns down and crosses into loss.
Both effects are positive when the hidden path is benign — the surrogate is a good guide. Push d below −b and you enter the paradox: surrogate up, truth down, every measurable correlation still reassuring. That gap is a year of budget aimed the wrong way.
How does this happen? The paradox arises when the treatment reaches the outcome through more than one path and those paths carry different signs, or when the surrogate–outcome association is confounded rather than causal. An ad-load increase might raise short-term revenue (one positive path) while eroding user trust and future engagement (a negative path the surrogate never sees). Because the surrogate captures only the first path, optimizing it maximizes the wrong sum. VanderWeele's contribution was to derive sufficient conditions under which a surrogate is "consistent" — provably free of the paradox — even when the treatment has a direct effect on the outcome not routed through the surrogate. Those conditions (monotonicity of effects, sign restrictions) are exactly the kind of structural assumptions that are hard to verify and easy to violate in messy media systems.
A parallel line, principal stratification (Frangakis and Rubin, 2002), reframes surrogacy around principal strata — groups defined by their joint potential surrogate values under treatment and control — and defines a "principal surrogate" through associative and dissociative effects. It is a cleaner causal object than Prentice's regression-flavored criterion, but it identifies less from data alone; principal surrogacy generally needs either multiple trials or strong structural assumptions to pin down. The practical takeaway from three decades of this work is consistent: a single surrogate almost never provably escapes the paradox, and the correlations you can see do not certify it. That negative result is what makes the next idea so useful — it changes the strategy from finding one perfect surrogate to combining many imperfect ones.
The Surrogate Index
The modern reframing comes from Athey, Chetty, Imbens, and Kang (2019; published in the Review of Economic Studies, 2025). Their move is deceptively simple. Stop searching for the one metric that fully mediates the long-run effect. Instead, collect many short-term proxies and combine them into a single predictor of the long-run outcome — the surrogate index — estimated on a historical sample where you do eventually observe the truth.
Definition: The surrogate index construction
Let \( S \) be a vector of short-term surrogates (many of them) and \( X \) baseline covariates. In an observational or historical sample where the long-run outcome \( Y^{L} \) was eventually observed, estimate the conditional mean
$$ h(s, x) \;=\; \mathbb{E}\!\left[\, Y^{L} \,\mid\, S = s,\; X = x \,\right]. $$This function \( h \) is the surrogate index: the predicted long-run outcome given everything you can see early. Now take a fresh experiment that ran only long enough to observe \( S \) and \( X \), and impute each unit's long-run outcome by \( h(S_i, X_i) \). The long-run average treatment effect is recovered as the effect on the index:
$$ \mathbb{E}\!\left[\, Y^{L}(1) - Y^{L}(0) \,\right] \;=\; \mathbb{E}_{\mathrm{exp}}\!\left[\, h(S, X) \mid W = 1 \,\right] - \mathbb{E}_{\mathrm{exp}}\!\left[\, h(S, X) \mid W = 0 \,\right]. $$You never wait for \( Y^{L} \) in the experiment; you borrow the surrogate-to-outcome map from history and apply it to the experiment's short-term data.
Two things make this more than a repackaging of Prentice. First, the crucial identifying condition is now stated on the whole vector of surrogates rather than a single marker. Using many proxies makes the required conditional independence far more plausible: it is easy to believe that some single metric leaves a path to the outcome open, and much harder to believe that a rich basket — early sales, search interest, brand-tracking scores, retention signals, engagement depth — jointly leaves any path uncovered. The assumption is the same in form; it is dramatically weaker in practice because a high-dimensional surrogate has more ways to capture the effect. A companion paper (Athey, Chetty, and Imbens, 2016) develops the "surrogate score" and surrogate index together and makes the many-surrogates logic explicit.
Second, the payoff is not only speed but precision. In their canonical application — a multi-site job-training experiment in California — the authors show that the first six quarters of short-run outcomes, indexed, recover the nine-year employment effect without the seven-year wait, and with materially tighter standard errors than the direct long-run estimate would have delivered. The surrogate index is thus a variance-reduction device as much as a time machine: a well-chosen index can be more efficient than observing the noisy long-run outcome directly, because the index averages over many informative early signals.
The Identifying Assumptions
None of this is free, and the two load-bearing assumptions deserve to be stated plainly, because violating either silently returns a confident wrong number.
💡 Surrogacy: the treatment reaches the long run only through the surrogates
The identifying condition is a conditional independence — the long-run outcome is independent of treatment given the surrogates and covariates:
$$ Y^{L} \;\perp\; W \;\mid\; S,\, X. $$This is Prentice's full-mediation criterion, promoted to a vector \( S \). It says the surrogates collectively capture every channel by which treatment moves the long-run outcome. The more (and more diverse) the surrogates, the more credible this is — but it is an assumption about causal structure, not something the data can fully confirm, and it fails exactly when there is a slow, hidden path the early metrics don't touch.
The second assumption is comparability (sometimes "stability" or "sample exchangeability"): the surrogate-to-outcome relationship \( h \) estimated in the historical sample must be the same relationship that holds in the experimental sample. If the map from early signals to long-run value drifted — because the product changed, the customer mix shifted, the macro environment turned, or the channel matured — then imputing with a stale \( h \) transports the wrong function. In media this is a serious hazard: the relationship between, say, early search lift and eventual revenue is not a law of nature but a regime that platforms, competitors, and seasons all perturb. The surrogate index inherits the external-validity burden of any transported estimate.
A subtle corollary: because \( h \) is estimated, it must be re-estimated. A surrogate index is not a formula you calibrate once and trust forever. It is a model of a relationship that decays, and keeping it honest requires periodically observing the true long-run outcome on some cohort and checking that the index still predicts it. This is the bridge to the industry practice that predates the econometrics.
Long-Term Holdbacks in Practice
Long before the surrogate index had a name, large platforms were solving the same problem operationally with long-term holdbacks. Hohnhold, O'Brien, and Tang (2015) is the canonical account. Studying ad-load changes on Google Search, they observed that the short-term revenue effect of showing more ads is a poor guide to the long-term effect, because users learn: served more ads, they grow "ad-blind" and click less over time; served better or fewer ads, they become "ad-sighted." The immediate metric and the durable metric can carry opposite signs.
Their instrument was a set of long-running holdback experiments — cohorts held out of a change for extended periods — that let them observe the user-learning curve directly and quantify how the effect evolves from launch toward a stable long-run level. The methodological point generalizes far beyond ad load: whenever behavior adapts to a treatment, the treatment's true effect is the converged effect, and you can only see convergence by keeping a control arm alive long enough to reach it. A long-term holdout is the empirical anchor that a surrogate index is trying to approximate without the wait.
💡 The two techniques are complements, not rivals
A long-term holdback observes the truth slowly and expensively but reliably; a surrogate index imputes it quickly and cheaply but on assumptions. The mature program runs both: use the occasionally-observed holdout truth to fit and re-validate the surrogate index, then use the index to make the many fast decisions in between. The holdback is the calibration; the index is the extrapolation. Neither is trustworthy alone — an index with no ground truth drifts undetectably, and a holdback with no index is too slow to steer week-to-week budgets.
Goodhart and the Discipline of Truth
There is a reason surrogate optimization deserves suspicion even when the statistics are clean: the moment a proxy becomes a target, the system starts to game it. This is Goodhart's law — "when a measure becomes a target, it ceases to be a good measure" — and it is the behavioral counterpart of the surrogate paradox. A team told to maximize short-term conversions will find ways to maximize short-term conversions, including ways that strip out long-run value the metric never counted. The optimization pressure itself degrades the surrogate's validity, and it does so precisely because the surrogate was chosen for its convenience rather than its completeness.
The defense is not to abandon proxies — you cannot; the long run is unobservable in real time — but to treat every surrogate as provisional and auditable. Three habits keep it honest. First, prefer an index to a single metric, so that no one dimension can be hill-climbed in isolation without moving the others. Second, keep observing the real outcome on a maintained holdout, and re-fit the surrogate model whenever the true outcome is available; a surrogate that has not been checked against truth recently is a liability dressed as a KPI. Third, watch for divergence: when the surrogate says "up" and the (slowly arriving) truth says "flat," treat that as a five-alarm signal that the map has broken or the paradox has arrived, not as noise to be smoothed. Proxy-hacking is not defeated by a cleverer proxy. It is defeated by refusing to let the proxy be the final word.
Surrogates and the MMM
Where does a marketing-mix model sit in this picture? An MMM already contains a small, honest model of persistence: adstock, the geometric or delayed decay that lets an impression this week keep paying into sales for several weeks after. Adstock is, in effect, a parametric surrogate for medium-run carryover — it extrapolates a short window of exposure into a slightly longer window of effect. For fast-and-medium channels that is often enough.
But genuine long-run brand effects operate on a horizon that ordinary adstock cannot reach. A typical MMM is fit on one to three years of weekly data with carryover windows measured in weeks; the brand equity that a sustained campaign builds compounds over years and across purchase cycles, well outside any adstock kernel the data can identify. Asking a standard MMM to see multi-year brand payback is asking it to extrapolate far beyond its window — precisely the regime where its uncertainty should balloon and its point estimates should not be trusted.
This is where surrogate thinking earns its place in the MMM stack. Rather than stretching adstock to implausible half-lives, bring in short-term surrogates that plausibly mediate the long-run effect and can be measured now: brand-tracking survey movements, branded-search interest, retention and repeat-purchase cohorts, consideration and awareness lifts. A dual-model approach — a short-term MMM for immediate response and a long-term or brand-equity layer that consumes these surrogates — separates the fast effect the weekly data can identify from the slow effect it cannot. The framework's own long-term brand modeling and its continuous-learning loop lean in exactly this direction: treat the medium-run response curve as directly estimable, and treat the genuinely long-run brand contribution as a surrogate-mediated quantity that must be calibrated against occasionally-observed truth rather than read off the adstock tail.
The connective tissue is triangulation. No single instrument sees the whole horizon: the MMM sees weeks, the lift test sees a campaign, the surrogate index projects the long run from early signals, and the long-term holdback eventually delivers ground truth. Each is biased in a different direction and validates the others. A surrogate index whose projections systematically disagree with the holdback is telling you its map has drifted; an MMM whose brand contribution disagrees with brand-tracking surrogates is telling you its long-run extrapolation is unmoored. Keeping all of them in the room, and keeping the real long-run outcome under measurement, is how you avoid mistaking a convenient proxy for the thing you actually sell against.
A Practical Synthesis
The discipline of long-run measurement comes down to a handful of rules that the theory and the industry precedent agree on. Choose surrogates that plausibly mediate the effect you care about — signals that sit causally between the spend and the long-run outcome, not merely correlated bystanders. Prefer an index over any single metric, because a rich basket makes the surrogacy assumption credible where a lone proxy makes it fragile and gameable. Estimate the surrogate-to-outcome map on data where you eventually saw the truth, and re-estimate it as the world moves, because that map is a regime, not a constant. Keep a long-term holdout alive so you always have fresh ground truth to calibrate against and to catch the surrogate paradox before it costs a year of misdirected budget. And never let a short-term proxy quietly become the objective without continually re-checking that it still tracks the outcome you are actually optimizing for. The long run is unobservable in real time — that is not going to change. What changes, with these tools, is whether your short-term decisions are aimed at a defensible estimate of the long run or at a convenient number that only looks like one.
Takeaways
- Budget decisions are bets on long-run value (LTV, retention, brand), but our instruments — lift tests, MMMs, platform metrics — mostly observe the short run. Optimizing a short-term proxy can actively push the wrong way.
- Prentice's criterion demands a valid surrogate fully mediate the treatment's effect on the true outcome (\( Y^{L} \perp W \mid S \)). A single surrogate almost never earns this, and the correlations you can see do not certify it.
- The surrogate paradox: a treatment can improve the surrogate, the surrogate can predict the outcome, and the treatment can still harm the outcome. Marginal association is no protection (VanderWeele; Frangakis & Rubin).
- The surrogate index (Athey, Chetty, Imbens, Kang) combines many short-term proxies into a predictor of the long-run outcome, fit on historical data and applied to a short experiment — making the surrogacy assumption far more plausible and often more precise than the direct estimate.
- Two assumptions carry the method: surrogacy (\( Y^{L} \perp W \mid S, X \)) and comparability between the historical and experimental samples. Both fail silently if a slow hidden path exists or the surrogate-to-outcome map has drifted.
- Long-term holdbacks (Hohnhold et al., Google) are the empirical anchor: use occasionally-observed truth to calibrate and re-validate the surrogate index, and to catch Goodhart-style proxy degradation.
- For MMMs: adstock is a medium-run surrogate; genuine multi-year brand effects exceed it. Bridge with mediating surrogates (brand tracking, branded search, retention cohorts), triangulate, and keep measuring the real long-run outcome.
References
- Prentice, R. L. (1989). Surrogate Endpoints in Clinical Trials: Definition and Operational Criteria. Statistics in Medicine, 8(4), 431–440.
- Athey, S., Chetty, R., Imbens, G. W., & Kang, H. (2019). The Surrogate Index: Combining Short-Term Proxies to Estimate Long-Term Treatment Effects More Rapidly and Precisely. NBER Working Paper 26463; published in the Review of Economic Studies (2025).
- Athey, S., Chetty, R., & Imbens, G. W. (2016). Estimating Treatment Effects Using Multiple Surrogates: The Role of the Surrogate Score and the Surrogate Index. arXiv:1603.09326.
- Hohnhold, H., O'Brien, D., & Tang, D. (2015). Focusing on the Long-term: It's Good for Users and Business. Proceedings of the 21st ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD), 1849–1858.
- VanderWeele, T. J. (2013). Surrogate Measures and Consistent Surrogates. Biometrics, 69(3), 561–565.
- Frangakis, C. E., & Rubin, D. B. (2002). Principal Stratification in Causal Inference. Biometrics, 58(1), 21–29.
- Prentice, R. L. (1989). See also Fleming, T. R., & DeMets, D. L. (1996). Surrogate End Points in Clinical Trials: Are We Being Misled? Annals of Internal Medicine, 125(7), 605–613.
- Goodhart, C. A. E. (1975). Problems of Monetary Management: The U.K. Experience. Papers in Monetary Economics, Reserve Bank of Australia. (Origin of "Goodhart's law.")
- Chetty, R. (2019). The Surrogate Index (Opportunity Insights replication data and lecture notes). Harvard University.