What changes in the ROAS / mROAS optimization if the outcome is customer lifetime value rather than sales, given that CLV is a model-based forecast observed with delay and censoring?

Summary

The optimization keeps its shape (maximize posterior-expected value under a budget, equalize marginal returns) but the numerator stops being an observed quantity. Incremental value becomes incremental customers × their model-based value + the change in value of existing customers, so a second posterior (the CLV model’s) multiplies the response-curve posterior, and margin, discount rate and the dropout tail become first-order inputs. Because realized lifetime value is right-censored for every recent cohort, a short-window outcome is only usable through a model that separates “not yet” from “never”; fixed windows or constant multipliers favour fast-payback channels. And because “acquired” is a post-treatment variable, the value of ad-acquired customers is identified only by an experiment that measures value per assigned unit in both arms, zeros included, never by comparing the customers each channel is credited with.

Answer

1. What stays and what changes

The sales version is in ROAS, mROAS, and Optimal Media Mix: ROAS is a counterfactual difference in predicted sales over the change period plus carryover periods, divided by spend (^roas-eq); the optimal mix maximizes summed predicted sales subject to (^optimal-mix); the static optimum is Dorfman–Steiner, margin × marginal response = 1 (^thm-dorfman-steiner). CLV is (^def-clv-decomposition), with at the individual level (^thm-det).

Synthesis: write for incremental acquisitions from channel (a Hill-type response with parameters ), for the mean model-based CLV of those customers (CLV parameters ), and for the change in value of the existing base. Then

Margin is inside CLV, so the optimum is unconstrained, or equal across channels under a budget.

Sales outcomeCLV outcome
NumeratorObserved sales, modelled counterfactuallyA forecast of discounted future margin
Horizon, with weeksInfinite, discounted at
Tail mechanismAdstock: memory of the adRepeat purchase: lifetime of the customer
UncertaintyPosterior of Joint posterior of plus margin, , model choice
SaturationHill curve on volume (Shape (Saturation) Effects)Hill on customer count and possibly falling marginal customer quality
Ground truthArrives within weeksRight-censored for years
Main failureExtrapolating the response curveAttributing value by credited channel; ranking channels on a truncated window

Two tails, easily double counted. The geometric lag with purchase feedback already warns that the lagged-sales coefficient “captures both advertising carryover and the repeat-purchase rate” and that the two must be separated (^def-glpf). A sales MMM with a 13-week window books the repeat purchases of newly acquired customers inside that window as media effect. Synthesis: under a CLV objective the MMM or experiment should model acquisitions (adstock still applies to them), the CLV model supplies everything after acquisition, and the two are multiplied, not added to sales-based ROAS.

The discount rate becomes first-order. On CDNOW the dropout heterogeneity parameter is . Synthesis: a Pareto II lifetime with has no finite mean, and the expected-purchases formula grows without bound in (^thm-pnbd-mean). Value is finite only through discounting (the factor in DET), so and margin (30% is simply assumed for CDNOW) move channel rankings as much as any response parameter, and neither is inside any posterior.

2. Acquisition versus retention effects

Acquisition. A new customer has no recency or frequency, so the individual-level machinery has nothing to condition on: gamma-gamma spend shrinks fully to the population mean at (^thm-gg-conditional-expectation), and BTYD “does not apply to new customers” in the sense of differentiating them (Bayesian and Hierarchical Extensions of CLV Models). Model choice bites hardest exactly here: for a customer with no repeat purchase Pareto/NBD gives on CDNOW while BG/NBD forces it to 1, and the zero class holds about 5% of cohort value. Differences between channels must therefore enter at the population level: acquisition channel as a time-invariant covariate, (^thm-clv-covariates), a cohort-indexed hierarchical fit, or a ZILN model on day-one features.

Illustrative ranking flip (own arithmetic, built on the vault's numbers)

Take the covariate example’s for paid social, a purchase-rate ratio of . Individual DET is linear in , so future value scales by roughly 0.67. With CDNOW’s $47 average customer value as the baseline: channel A at $30 per incremental customer returns ; paid social at $25 returns . On first-purchase ROAS paid social wins (same basket, lower cost); on CLV it loses.

Retention. The BTYD and sBG models assume constant individual traits, admit only time-invariant covariates, and their authors describe forecasts as “a baseline against which we can examine the impact of changes in marketing activity”, warning that marketing covariates invite “endogeneity bias and sample selection bias” when targeting used past RFM. So cannot be read from a fitted CLV model. It needs randomization among existing customers, after which arm is a legitimate covariate and the estimand is the arm difference in DET. Synthesis, a trap: an ad that merely pulls purchases forward raises frequency and recency inside the window, so the posterior on and , and hence forecast CLV, rises although nothing durable changed. The model reads a transient as a trait because stationarity is its core assumption. Check this by comparing arms again at a later read-out.

3. Are ad-acquired customers different? Selection at three levels

  1. Marginal versus average. is the quality of the customers added by the next dollar. The vault’s Hill curves saturate volume; nothing in it models declining quality, but the term is in the derivative and is plausibly negative.
  2. Credited versus caused. Activity Bias in Advertising shows exposed users are more active regardless of ads. Customers credited to a channel are therefore enriched with would-have-bought-anyway, high-RFM people, inflating that channel’s observed mean CLV. The covariate note says it directly: whether a channel “caused lower-value customers or merely selected them is not identified by this model”.
  3. Post-treatment conditioning. Synthesis: “became a customer” is a post-treatment variable, so comparing acquired customers across arms compares different principal strata, always-buyers versus always-buyers plus ad-induced buyers (Instrumental Variables and Principal Stratification). The identified quantity is the intention-to-treat effect on value per assigned unit, non-customers counted as zero. The mean value of ad-induced customers is the complier ratio , valid only under monotonicity and the exclusion restriction that ads do not change always-buyers’ value, which is precisely “no retention effect”. The allocation needs only the total ITT; the acquisition/retention split is a further assumption.

A time-series version of the same trap is the ruse of heterogeneity: aggregate retention rises with tenure though no individual changes (^thm-ruse-of-heterogeneity). A channel scaled recently has young cohorts and looks worse on blended retention. Compare cohorts at equal tenure, and distrust linear RFM scores, which miss the backward-bending iso-value curves (^thm-increasing-frequency-paradox): a promotion-driven burst followed by silence signals low value.

4. Delay and censoring: four kinds of lateness

LatenessExampleSame idea asRepair
Effect delayAdstockDistributed lagWindow of ; see Q - How Adstock Breaks Switchback and Sequential Test Assumptions
Customer tailRepeat purchases for yearsPurchase feedbackTransaction-flow model
Measurement delayConversion 30 days after clickRight censoring with a “never” classJoint “ever” and “when” likelihood
Administrative censoringYoung cohort, short historysBG censoring term, Kaplan–MeierCohort likelihood with a survivor term; pooling across cohorts

Genuinely the same idea: the delayed feedback model’s weight that a silent click will still convert (^def-em-estep) and Pareto/NBD’s , which decays in the silence at rate . Both are posteriors over “not yet versus never”, and both extend Survival Analysis with an event that may never occur. Survival analysis also requires censoring to be noninformative: cohort age is administrative and harmless in itself, but if the channel mix shifted over time, the youngest and most model-dependent cohorts belong disproportionately to the newest channels.

Three lessons transfer to CLV:

  • The window dilemma is fatal here. Short windows mislabel, long windows are stale (Chapelle: 30 days stale, with new campaigns already 11.3% of traffic after 26 days). A “true” CLV label needs years, so waiting is not an option.
  • Constant multipliers are the Rescale baseline. Scaling 90-day revenue by one LTV multiplier assumes every cohort matures alike. Chapelle’s analogue assumed labels missing at random and underpredicted by 30% on recent campaigns, because missingness depends on elapsed time.
  • A hard window reorders channels. With a window the delayed bandit equals an immediate one with rates (^thm-bandit-lb-censored), harmless for ranking only because the delay CDF is assumed shared across arms. Synthesis: channels differ in payback curves, so a windowed-value optimizer systematically favours fast-payback channels.

Model-based CLV is thus the vault’s version of a surrogate: a function of short-window that, if the model holds, uses all the information in that window (RFM sufficiency). Its weak point is identification from young cohorts, the CLV counterpart of the delayed-feedback model’s two basins (low rate and short delay versus high rate and long delay). The sBG result shows structure pays: seven years of data project year-12 survival within about 4%, where curve-fitting regressions miss by 30 to 92%.

5. Propagating the CLV posterior into the decision

Decision Analysis gives and From Inference to Decision says to “propagate it as necessary to any decisions” (^def-expected-value-rule). Here :

  • Plug in joint draws, never posterior means (^posterior-plug-in). is a ratio with a heavy right tail, so CLV at mean parameters is not mean CLV. PyMC-Marketing returns per-customer draws; note that its CLV is a finite-horizon sum, not the closed-form infinite-horizon DET.
  • Know where the uncertainty is. With tens of thousands of customers population parameters are tight; uncertainty concentrates in small or young cohorts and in covariate coefficients, which is exactly the channel-quality term. Pool cohorts hierarchically rather than fitting each alone.
  • Optimize the expectation, report the distribution. Route A (optimize average value over draws) yields a stable mix; route B (optimize per draw) shows how uncertain it is. The sales-only optimum was already multimodal when extrapolating; multiplying by widens it.
  • The expected-value rule has conditions: many small independent decisions, no foreclosed options, a trusted model. Bid-level choices satisfy the first; an annual budget split does not. Conversely, do not demand 97.5% certainty that one channel beats another before moving money; for partially pooled estimates that threshold is “very strict”.
  • Score the forecast properly. CLV is zero-inflated and heavy-tailed (the top CDNOW cell, 954 customers, holds 38% of value). CRPS is in dollars and reduces to absolute error for point forecasts, so BTYD draws, ZILN and a naive multiplier are comparable (^def-crps). Backtest by rolling origin over acquisition cohorts, avoid MAPE with zeros, and check coverage of cohort sums, since marginals can be calibrated while sums are not (Forecast Evaluation and Backtesting). Fader and Hardie’s tracking plot and conditional expectations are the same checks; run them by channel and by arm.

Practical Implications

What an experiment must measure.

  1. Randomized exposure (user-level or geo) with identity resolution, so customers and their transactions map to an arm.
  2. Outcomes for all assigned units: value per capita or per geo, zeros included. Never only converters.
  3. New-customer counts by arm, and a transaction log (id, date, amount) long enough to contain repeat purchases, giving per arm.
  4. Existing customers’ transactions by arm, for .
  5. Spend by arm, so that is computed by dividing posterior draws, as TBR does for iROAS (^def-iroas).
  6. Time-to-purchase distributions by arm, to separate incrementality from acceleration.
  7. Pre-registered margin, , horizon, model family and zero-class treatment, plus scheduled re-reads (for example 3, 6 and 12 months) in which forecasts are scored against realized cohort revenue.

Decision rule. Model acquisitions in the MMM or geo test; multiply by an arm- or channel-specific CLV posterior fitted on cohorts, not on credited customers; optimize expected value over joint draws; report the distribution of the optimum; show sensitivity to , margin and Pareto/NBD versus BG/NBD before acting. If the CLV ranking differs from the sales ranking only through a covariate estimated on observational channel labels, treat it as a hypothesis for the next experiment.

For the ABM. Give agents latent drawn from gamma mixing distributions by acquisition source, as Heterogeneity in Agent Models recommends, and calibrate to conditional expectations by frequency class. A homogeneous loyalty-growth rule would reproduce rising retention for the wrong reason.

Source Notes

NoteRelevance
Customer Lifetime Value - OverviewCLV decomposition, baseline-not-causal warning, validation standard
Pareto-NBD Model · BG-NBD Model, new-customer expectation, zero-class difference,
Gamma-Gamma Model of Monetary ValueShrinkage of spend; full shrinkage at ; independence assumption
RFM Sufficient Statistics and Iso-Value CurvesDET, discounting, increasing-frequency paradox, CDNOW cohort values
Shifted-Beta-Geometric Model for Contractual RetentionRuse of heterogeneity, censored cohort likelihood, projection accuracy
Bayesian and Hierarchical Extensions of CLV ModelsCovariates, full Bayes, cohort pooling, ZILN, channel-covariate example
ROAS, mROAS, and Optimal Media Mix · Optimal Marketing Decisions and Forecasting · Shape (Saturation) EffectsSales-based objective, posterior plug-in rule, optimum instability, Dorfman–Steiner, Hill curve
Carryover Effects and Distributed LagsPurchase feedback versus advertising carryover
Delayed and Censored Feedback - Overview · Delayed Feedback Model for Conversion Prediction · EM and Gradient Optimization for the Delayed Feedback Model · Bandit Models with Delayed and Censored FeedbackWindow dilemma, “ever” and “when”, Rescale failure, censored bandit equivalence
Survival AnalysisRight censoring; noninformative-censoring requirement
Decision Analysis · From Inference to DecisionExpected-utility rule and its three conditions; against thresholds
Proper Scoring Rules (CRPS, Log Score, Pinball Loss) · Forecast Evaluation and BacktestingCRPS, rolling origin, coverage of sums
Instrumental Variables and Principal Stratification · Activity Bias in AdvertisingPost-treatment strata; selection of credited customers
Fader Hardie Lee 2005 - RFM and CLV Iso-Value CurvesSecs. 1-5, Eq. 2, Tables 2-3

Gaps

  • No note on surrogate indices or long-term effect estimation from short-term outcomes. The surrogacy reading of model-based CLV above is my synthesis; “surrogate endpoints” appears in the vault only as a one-line application of principal stratification.
  • No model of marginal customer quality as a function of spend, and no acquisition-versus-retention budget allocation literature.
  • No CLV model with time-varying marketing covariates or non-stationary traits; the ingested models exclude the retention effect by construction. Abe (2009) is cited from memory in the extensions note, not ingested.
  • No value-of-information treatment of how long to wait before re-allocating, and no empirical evidence in the vault that ad-acquired customers differ in CLV; the example is illustrative.

Follow-Up Questions

  • How should a hierarchical BG/NBD with arm and acquisition-cohort effects be specified so that the ITT on value per assigned unit comes with a posterior?
  • How sensitive is the optimal mix to when , and should the discount rate itself carry a prior?
  • What read-out schedule minimizes expected loss from acting on an immature CLV forecast versus waiting?