RFM Sufficient Statistics and Iso-Value Curves
Summary
Direct marketers have long scored customers on recency, frequency and monetary value (RFM). Fader, Hardie & Lee (2005, JMR) show that under a Pareto-NBD Model for transactions and an independent Gamma-Gamma Model of Monetary Value for spend, RFM are not merely good predictors — they are sufficient statistics for a customer’s purchase history. The paper derives a closed form for discounted expected transactions (DET) over an infinite horizon, multiplies it by shrunken expected spend and margin to get CLV, and visualizes the result as iso-value curves in the recency–frequency plane. The curves bend backwards: among customers who have not bought recently, more past purchases imply lower value — the increasing frequency paradox. Applied to 23,560 CDNOW customers, the cohort is worth about $1.1 million, with roughly 5% of it sitting in the “zero class” of customers who never repurchased.
Overview
An RFM scoring model regresses period-2 behaviour on period-1 R, F and M. The paper’s objection (Customer Lifetime Value - Overview) is that such a model predicts one period ahead, burns half the data on a dependent variable, and treats noisy summaries as if they were the latent traits. The alternative: keep RFM as the only inputs, but let a behavioural model determine how they map to future value. The exploratory plots make the case. Average holdout spend by for the full cohort is “very sparse and therefore somewhat untrustworthy” despite 23,560 customers; its contour plot is jagged; and its shape depends on the arbitrary lengths of the two periods. A model “fills in the blanks” and removes the dependence on the observation window.
Main Content
RFM as sufficient statistics ^def-rfm-sufficient
Notation : = number of repeat transactions in (frequency), = time of the last one (recency; if ), = time since acquisition. With = average value of the transactions (monetary value):
- The Pareto/NBD individual likelihood depends on the purchase times only through (likelihood) because Poisson purchasing is memoryless.
- The gamma-gamma likelihood depends on the individual amounts only through because the gamma is closed under convolution.
Hence “we formally link the observed measures to the latent traits and show that no other information about customer behavior is required in order to implement our model” (Sec. 1). Note that “recency” here is the time of the last purchase measured from acquisition — larger is more recent.
Discounted expected transactions under Pareto/NBD ^thm-det
A discrete-time sum forces arbitrary choices of horizon and period length, and thresholding to pick a horizon “goes against the spirit of the model.” Instead, work in continuous time over an infinite horizon. Conditional on the traits, the present value of the transaction stream is (Eq. A3)
Multiplying by and averaging over the posterior of gives (Eq. 2 / Appendix Eqs. A8–A10)
where is the confluent hypergeometric function of the second kind (Tricomi’s ) and is the NBD likelihood. The continuously compounded rate for data in periods per year is ; a 15% annual rate with weekly data gives . The authors call this “a new analytical result … central to our CLV estimation.”
CLV from RFM ^def-clv-from-rfm
Substituting DET and the gamma-gamma conditional expectation (Eq. 4) into Eq. 1:
Seven population parameters , a margin and a discount rate turn any triple into a lifetime value. An iso-value curve is a level set of this function in the plane for fixed and .
Using the customer’s raw instead of “would ignore the ‘regression-to-the-mean’ phenomenon”; the shrunken value is the correct DET multiplier. For it equals the population mean ($36 on the 78-week data).
The increasing frequency paradox ^thm-increasing-frequency-paradox
Except at , DET increases in recency, with a strong interaction: “for low frequency customers, there is an almost linear relationship between recency and DET”, but it is “highly nonlinear for high frequency customers” (Sec. 4). In low-value regions the iso-value lines bend backwards: a customer with , has DET , “the same as someone with a lower (i.e., worse) frequency of and recency of . In general, for people with low recency, higher frequency seems to be a bad thing.”
Explanation (Fig. 11). If both customers were known to be alive, the frequent buyer would be worth more. But a long silence is far less likely under a high purchase rate, so the posterior puts most of its mass on “dead”. In the paper’s illustration, the sparse-but-steady customer A has DET 4.6 versus 1.9 for customer B, whose nine purchases all came early. Formally, decays in at rate (P(alive)).
The authors draw the methodological moral: “The use of a regression-based specification, which is used in many scoring models, would likely miss this pattern and lead to faulty inferences for a large portion of the recency-frequency space.” A linear-in-RFM score is monotone in frequency by construction.
Validation before valuation (Sec. 3)
On the 2,357-customer sample (39 weeks calibration / 39 holdout): Pareto/NBD MLEs track cumulative repeat sales with under 2% error at week 78, and conditional expectations by calibration frequency follow the actuals. Gamma-gamma MLEs are . A combined test multiplies expected spend per transaction by (a) actual and (b) predicted holdout transactions: (a) isolates the spend model and shows no bias that would contradict the independence or stationarity assumptions; (b) tests the full system. Re-estimating on all 78 weeks barely changes the fit (39-week Pareto/NBD log-likelihood at the 78-week parameters versus at the optimum), which justifies using all the data for the final valuation — something a two-period scoring model cannot do.
Valuing the CDNOW cohort (Sec. 4, Tables 2–3)
Customers with repeat purchases (11,506 of 23,560) are coded into terciles on each of R, F and M, giving 27 cells plus the zero cell (12,054 customers, R=F=M=0). Margin 30%, discount rate 15%.
- Zero class. Average CLV about $4.40 each, but $53,000 in total — “almost 5% of the total future value of the entire cohort — larger than most of the 27 other RFM cells.” “This slight whisper of CLV becomes a loud roar when applied to such a large group.”
- Top cell (R=F=M=3). 954 customers, $414,900 in total, about $435 each — nearly 38% of cohort value.
- Whole cohort. About $47 per customer, just over $1.1 million.
- Within each M level, high-frequency/low-recency cells are worth less than lower-frequency ones (e.g. M=3, R=1: $11,300 across 676 customers for F=1, about $17 each, versus $1,000 across 101 customers for F=3, about $10 each) — the paradox in tabular form.
| Average CLV by tercile | 1 | 2 | 3 |
|---|---|---|---|
| Recency | $10 | $62 | $201 |
| Frequency | $18 | $50 | $205 |
| Monetary value | $31 | $81 | $160 |
Recency discriminates most and monetary value least, “consistent with the widely-held view that recency is usually a more powerful discriminator” — hence “RFM” rather than “FRM”.
Limitations stated by the authors (Sec. 5): no marketing-mix covariates (and adding them risks endogeneity when targeting used past RFM); independence between spend and transactions; constant margin; noncontractual setting only; a single cohort — managers “need to run the model across multiple cohorts”. They suggest a bivariate Sarmanov distribution or a hierarchical Bayesian formulation to correlate and .
Examples
Tracing a backward-bending curve. Using the 39-week Pareto/NBD estimates, and (own calculation; the paper’s Figure 10 uses unreported 78-week re-estimates, so levels differ slightly):
| DET | Comment | |
|---|---|---|
| 0.18 | zero class | |
| 1.43 | one purchase, long ago | |
| 0.89 | more purchases, slightly more recent, lower value | |
| 1.07 | still below the single-purchase customer | |
| 3.12 | one purchase last week | |
| 14.07 | frequent and recent | |
| 29.19 | back corner of the surface |
With m_x=\50x=7E(M\mid m_x,x)=$49.1(7,70)0.30\times49.1\times14.07\approx$207(7,35) customer is worth about \13.
import numpy as np
from scipy.special import gammaln, hyperu
def pnbd_det(r, a, s, b, x, tx, T, delta, pnbd_lik):
"""Eq. 2. pnbd_lik(x, tx, T) returns the Pareto/NBD likelihood (not its log)."""
num = (a**r * b**s * delta**(s-1) * np.exp(gammaln(r+x+1) - gammaln(r))
* hyperu(s, s, delta*(b+T)))
return num / ((a+T)**(r+x+1) * pnbd_lik(x, tx, T))
delta = np.log(1.15) / 52 # 15% annual, weekly dataConnections
- Pareto-NBD Model — supplies the likelihood and inside DET.
- Gamma-Gamma Model of Monetary Value — supplies the DET multiplier and the (tested) independence assumption.
- Customer Lifetime Value - Overview — the CLV decomposition and the critique of scoring models.
- BG-NBD Model — an alternative transaction engine; its zero class has , so zero-class value would be treated differently.
- Empirical Bayes - Overview — the iso-value surface is a map of posterior expectations under a fitted prior; “filling in the blanks” is smoothing by partial pooling.
- Optimal Marketing Decisions and Forecasting — CLV by RFM cell is the input to retention and reactivation budget decisions.
See Also
- Markets Data and Sales Drivers — customer transaction databases as a marketing data source.
- Metalearners for CATE — RFM-based targeting needs incremental effects, not just value; CLV ranks customers by worth, CATE by responsiveness.
- Bayesian and Hierarchical Extensions of CLV Models — posterior uncertainty on CLV, covariates, and the ML alternative for brand-new customers with no RFM history.
- Survival Analysis — continuous-time discounting of a survivor function, , is the general CLV definition the appendix starts from.