RFM Sufficient Statistics and Iso-Value Curves

Summary

Direct marketers have long scored customers on recency, frequency and monetary value (RFM). Fader, Hardie & Lee (2005, JMR) show that under a Pareto-NBD Model for transactions and an independent Gamma-Gamma Model of Monetary Value for spend, RFM are not merely good predictors — they are sufficient statistics for a customer’s purchase history. The paper derives a closed form for discounted expected transactions (DET) over an infinite horizon, multiplies it by shrunken expected spend and margin to get CLV, and visualizes the result as iso-value curves in the recency–frequency plane. The curves bend backwards: among customers who have not bought recently, more past purchases imply lower value — the increasing frequency paradox. Applied to 23,560 CDNOW customers, the cohort is worth about $1.1 million, with roughly 5% of it sitting in the “zero class” of customers who never repurchased.

Overview

An RFM scoring model regresses period-2 behaviour on period-1 R, F and M. The paper’s objection (Customer Lifetime Value - Overview) is that such a model predicts one period ahead, burns half the data on a dependent variable, and treats noisy summaries as if they were the latent traits. The alternative: keep RFM as the only inputs, but let a behavioural model determine how they map to future value. The exploratory plots make the case. Average holdout spend by for the full cohort is “very sparse and therefore somewhat untrustworthy” despite 23,560 customers; its contour plot is jagged; and its shape depends on the arbitrary lengths of the two periods. A model “fills in the blanks” and removes the dependence on the observation window.

Main Content

RFM as sufficient statistics ^def-rfm-sufficient

Notation : = number of repeat transactions in (frequency), = time of the last one (recency; if ), = time since acquisition. With = average value of the transactions (monetary value):

  • The Pareto/NBD individual likelihood depends on the purchase times only through (likelihood) because Poisson purchasing is memoryless.
  • The gamma-gamma likelihood depends on the individual amounts only through because the gamma is closed under convolution.

Hence “we formally link the observed measures to the latent traits and show that no other information about customer behavior is required in order to implement our model” (Sec. 1). Note that “recency” here is the time of the last purchase measured from acquisition — larger is more recent.

Discounted expected transactions under Pareto/NBD ^thm-det

A discrete-time sum forces arbitrary choices of horizon and period length, and thresholding to pick a horizon “goes against the spirit of the model.” Instead, work in continuous time over an infinite horizon. Conditional on the traits, the present value of the transaction stream is (Eq. A3)

Multiplying by and averaging over the posterior of gives (Eq. 2 / Appendix Eqs. A8–A10)

where is the confluent hypergeometric function of the second kind (Tricomi’s ) and is the NBD likelihood. The continuously compounded rate for data in periods per year is ; a 15% annual rate with weekly data gives . The authors call this “a new analytical result … central to our CLV estimation.”

CLV from RFM ^def-clv-from-rfm

Substituting DET and the gamma-gamma conditional expectation (Eq. 4) into Eq. 1:

Seven population parameters , a margin and a discount rate turn any triple into a lifetime value. An iso-value curve is a level set of this function in the plane for fixed and .

Using the customer’s raw instead of “would ignore the ‘regression-to-the-mean’ phenomenon”; the shrunken value is the correct DET multiplier. For it equals the population mean ($36 on the 78-week data).

The increasing frequency paradox ^thm-increasing-frequency-paradox

Except at , DET increases in recency, with a strong interaction: “for low frequency customers, there is an almost linear relationship between recency and DET”, but it is “highly nonlinear for high frequency customers” (Sec. 4). In low-value regions the iso-value lines bend backwards: a customer with , has DET , “the same as someone with a lower (i.e., worse) frequency of and recency of . In general, for people with low recency, higher frequency seems to be a bad thing.”

Explanation (Fig. 11). If both customers were known to be alive, the frequent buyer would be worth more. But a long silence is far less likely under a high purchase rate, so the posterior puts most of its mass on “dead”. In the paper’s illustration, the sparse-but-steady customer A has DET 4.6 versus 1.9 for customer B, whose nine purchases all came early. Formally, decays in at rate (P(alive)).

The authors draw the methodological moral: “The use of a regression-based specification, which is used in many scoring models, would likely miss this pattern and lead to faulty inferences for a large portion of the recency-frequency space.” A linear-in-RFM score is monotone in frequency by construction.

Validation before valuation (Sec. 3)

On the 2,357-customer sample (39 weeks calibration / 39 holdout): Pareto/NBD MLEs track cumulative repeat sales with under 2% error at week 78, and conditional expectations by calibration frequency follow the actuals. Gamma-gamma MLEs are . A combined test multiplies expected spend per transaction by (a) actual and (b) predicted holdout transactions: (a) isolates the spend model and shows no bias that would contradict the independence or stationarity assumptions; (b) tests the full system. Re-estimating on all 78 weeks barely changes the fit (39-week Pareto/NBD log-likelihood at the 78-week parameters versus at the optimum), which justifies using all the data for the final valuation — something a two-period scoring model cannot do.

Valuing the CDNOW cohort (Sec. 4, Tables 2–3)

Customers with repeat purchases (11,506 of 23,560) are coded into terciles on each of R, F and M, giving 27 cells plus the zero cell (12,054 customers, R=F=M=0). Margin 30%, discount rate 15%.

  • Zero class. Average CLV about $4.40 each, but $53,000 in total — “almost 5% of the total future value of the entire cohort — larger than most of the 27 other RFM cells.” “This slight whisper of CLV becomes a loud roar when applied to such a large group.”
  • Top cell (R=F=M=3). 954 customers, $414,900 in total, about $435 each — nearly 38% of cohort value.
  • Whole cohort. About $47 per customer, just over $1.1 million.
  • Within each M level, high-frequency/low-recency cells are worth less than lower-frequency ones (e.g. M=3, R=1: $11,300 across 676 customers for F=1, about $17 each, versus $1,000 across 101 customers for F=3, about $10 each) — the paradox in tabular form.
Average CLV by tercile123
Recency$10$62$201
Frequency$18$50$205
Monetary value$31$81$160

Recency discriminates most and monetary value least, “consistent with the widely-held view that recency is usually a more powerful discriminator” — hence “RFM” rather than “FRM”.

Limitations stated by the authors (Sec. 5): no marketing-mix covariates (and adding them risks endogeneity when targeting used past RFM); independence between spend and transactions; constant margin; noncontractual setting only; a single cohort — managers “need to run the model across multiple cohorts”. They suggest a bivariate Sarmanov distribution or a hierarchical Bayesian formulation to correlate and .

Examples

Tracing a backward-bending curve. Using the 39-week Pareto/NBD estimates, and (own calculation; the paper’s Figure 10 uses unreported 78-week re-estimates, so levels differ slightly):

DETComment
0.18zero class
1.43one purchase, long ago
0.89more purchases, slightly more recent, lower value
1.07still below the single-purchase customer
3.12one purchase last week
14.07frequent and recent
29.19back corner of the surface

With m_x=\50x=7E(M\mid m_x,x)=$49.1(7,70)0.30\times49.1\times14.07\approx$207(7,35) customer is worth about \13.

import numpy as np
from scipy.special import gammaln, hyperu
 
def pnbd_det(r, a, s, b, x, tx, T, delta, pnbd_lik):
    """Eq. 2. pnbd_lik(x, tx, T) returns the Pareto/NBD likelihood (not its log)."""
    num = (a**r * b**s * delta**(s-1) * np.exp(gammaln(r+x+1) - gammaln(r))
           * hyperu(s, s, delta*(b+T)))
    return num / ((a+T)**(r+x+1) * pnbd_lik(x, tx, T))
 
delta = np.log(1.15) / 52          # 15% annual, weekly data

Connections

See Also