Interference and Marketplace Experiments
Summary
The difference-in-means estimator is unbiased for the average treatment effect only under SUTVA: a unit’s outcome may not depend on anyone else’s assignment. In social networks (treated users message control friends) and marketplaces (treated and control units compete for the same drivers, listings or ad slots) this fails, and the estimand of interest becomes the global treatment effect (GTE): everyone treated versus no-one treated. Johari, Li, Liskovich & Weintraub (2020; Management Science 2022) model a two-sided booking platform as a Markov chain with a mean-field ODE limit and show that the bias of the two standard designs depends on market balance: customer-side randomization (CR) is unbiased when demand-constrained and biased when supply-constrained; listing-side randomization (LR) is the reverse; both overestimate a positive effect. A two-sided randomization (TSR) design interpolates between them. Design-side alternatives are cluster randomization (networks, geos) and switchbacks (time), all of which buy lower bias with higher variance.
Overview
Larsen et al. (Sec. 6) give two canonical examples. LinkedIn messaging: a feature that makes treated users send more messages also makes their control-group friends reply more, contaminating control (network interference). Lyft pricing: a treatment that makes riders book more rides depletes the shared pool of drivers, lowering bookings in control (marketplace interference). In both cases “traditional randomization no longer adequately approximates the counterfactuals”. Reported magnitudes are large: Blake & Coey (2014) found an eBay auction experiment off by a factor of two; Fradkin (2019) a 50% overestimate in simulation; Holtz et al. (2020) estimated the interference bias on Airbnb at almost one third of the treatment effect.
There are two strategies (Larsen Sec. 6): design the interference away, by assigning units that influence each other to the same arm so that a difference in means is again meaningful, or model it, keeping unit-level randomization and its power but relying on a correctly specified interference model. Johari et al. sit in between: a structural market model is used to understand when simple designs are biased and to motivate a new design.
Main Content
SUTVA, interference, and the global treatment effect ^def-gte
With assignment vector , unit ‘s potential outcome is in general . SUTVA asserts . Interference (spillover, leakage) is any violation. The decision-relevant estimand is the global treatment effect
the difference in the (steady-state) outcome rate between a fully treated and a fully controlled market. Under SUTVA this equals the usual ATE; under interference a 50/50 experiment observes neither world.
The market model (Johari et al. Secs. 3-4)
listings of types (mass ) are either available or occupied. Customers of types arrive as Poisson processes with total rate per listing. An arriving customer includes each available type- listing in her consideration set with probability and chooses by multinomial logit with utilities and outside option . A booked listing stays occupied for an exponential time with rate . As the scaled state , the mass of available listings, follows the mean-field ODE (Eqs. 7-8)
Theorem 1 shows a unique, globally asymptotically stable steady state (via a convex Lyapunov function); Theorem 2 (Kurtz) shows the finite Markov chain converges to this fluid limit. The ratio is market balance: small means demand-constrained (few customers, inventory replenishes quickly), large means supply-constrained. A treatment is a change in choice parameters , e.g. better photos, badges, or showing completion rates; it is encoded by doubling the type space into control and treated copies.
Designs and naive estimators (Sec. 5)
Let be the booking rate over of customers in condition booking listings in condition .
- Customer-side randomization (CR): a fraction of customers is treated; .
- Listing-side randomization (LR): a fraction of listings is treated and every customer sees a mix; .
- Two-sided randomization (TSR): both sides are randomized and the intervention is applied only when a treated customer views a treated listing. The naive estimator (Eq. 21) is
which reduces to CR as and to LR as . A multiple-randomization design of this kind was proposed independently by Bajari et al. (2019).
Bias depends on market balance (Theorems 3-4, Proposition 4) ^thm-market-balance
In the mean-field steady state:
- Demand-constrained, : for all , whereas generically .
- Supply-constrained, : and , whereas generically the CR estimator stays biased.
- For a positive treatment ( everywhere) the biases in (1) and (2) are strictly positive: the naive estimators overestimate the GTE.
Intuition. When inventory replenishes between arrivals, customers never compete, so CR has no interference; but under LR each customer compares treated with control listings side by side, and the treated listings cannibalise bookings from control listings, so the contrast is inflated. When supply is scarce, every available listing gets booked anyway, so listings do not compete and LR is clean; but under CR treated customers book inventory that control customers would otherwise have found. In each biased case, “individuals in the treatment group face less competition than they would in the global treatment setting, whereas the individuals in the control group face more competition than in the global control setting”. In the supply-constrained limit the GTE itself vanishes (inventory, not demand, binds), so the LR relative bias need not vanish even though its absolute bias does.
Tuning TSR (Sec. 6.3). Choose allocations as a function of observable market balance (Eq. 26):
so TSR becomes CR when demand-constrained and LR when supply-constrained; Corollary 1 shows TSRN is unbiased in both limits. Because a TSR experiment observes all four cells , it measures competition directly. The heuristic estimators TSRI-1 and TSRI-2 start from an interpolation and subtract cross-cell correction terms weighted by a factor ; TSRI-2 had the lowest bias of all five estimators at intermediate balance.
Bias-variance trade-off (Sec. 7). In simulations with listings over 500 runs, the ordering of bias matches the mean-field theory, but the TSR estimators with the lowest bias have the highest variance; TSRN has variance similar to the better of CR and LR. Bias is insensitive to market size and horizon while variance shrinks with both, so large, long experiments should prioritise bias reduction and small, short ones variance.
Cluster randomization
For network interference the dominant design is graph-cluster randomization (Ugander et al. 2013; Eckles et al. 2014; Saveski et al. 2017): partition the graph by community detection so that most edges are within clusters, and randomize clusters. Units then mostly share treatment with their neighbours, approximating the all-treated and all-control worlds, but the effective sample size is the number of clusters and power drops sharply. Ego-cluster designs (Saint-Jacques et al. 2019) use many small clusters of one ego plus some alters, recovering power and allowing the spillover itself to be estimated by treating ego and alters differently. Inference must respect the randomization unit; see Standard Errors and Clustering. In Johari et al.’s comparison (Sec. 8), a cluster-randomized estimator beats TSR when the market is tightly clustered (customers strongly prefer one listing type) and loses when the market is interconnected, where no clean partition exists. For auctions, budget-split designs (Liu et al. 2021) give each arm its own copy of the budget so the arms no longer compete for it.
Geo experiments are cluster-randomized designs. Randomizing DMAs rather than users is precisely a response to interference (shared auctions, cross-device identity, offline sales) plus measurement constraints, and it pays the same price in effective sample size; see Geo-Experiment Design and Power Analysis. Residual interference appears as cross-border spillover.
Examples
Badge experiment on a lodging platform. The platform tests a “top host” badge that raises a listing’s utility from to (the paper’s Figure 2 parameters: a 20% steady-state booking probability under global control and 23% under global treatment at ).
- In low season (), randomize customers. An LR test would show badged listings far outperforming unbadged ones mainly because they divert bookings, not because total bookings rise.
- In peak season () almost everything books regardless; randomize listings. A CR test would show treated customers booking more only because they got to scarce inventory first.
- In between, run TSR with allocations from Eq. 26, or cluster by destination if travellers rarely substitute across destinations.
Advertising analogue. A bidding-algorithm test that splits campaigns (the supply of ads competing for impressions) is an LR-type design: the treated campaigns win auctions from the control campaigns and the measured lift overstates the global effect. Splitting users is CR-type and is biased when shared budgets or frequency caps bind. Budget-split, geo-cluster or switchback designs are the remedies.
Simulation as a design tool. The Markov-chain market model is a compact agent-based simulator: heterogeneous agents, a choice rule and inventory dynamics. Simulating global treatment, global control and each candidate design gives the bias and variance of every estimator before running anything live. The same approach transfers to richer agent-based market models, where the mean-field limit plays the role of an analytical check.
Connections
- Potential Outcomes Framework — SUTVA is part of the definition of ; interference requires .
- Switchback Experiment Design and Analysis — randomize the whole market over time instead of units within it.
- Geo-Experiment Design and Power Analysis and Geo-Experiment Methodology - Overview — spatial cluster randomization in marketing measurement.
- Standard Errors and Clustering — inference when the randomization unit is a cluster.
- Sample Ratio Mismatch and Trustworthiness Checks — a different failure: there the analysed sample is distorted, here the potential outcomes are.
- Online Experimentation - Overview — interference as one of the four failure modes of the naive A/B test.
- Observational vs Experimental Methods in Advertising — randomization alone does not guarantee the right estimand in ad markets.
- Differences-in-Differences and Synthetic Control — aggregate-unit alternatives when only a few markets can be treated.
See Also
- The Experimental Ideal
- Activity Bias in Advertising
- Randomization Inference - Overview
- Q - Carryover Dynamics and the Timing of Sequential Media Experiments
- Q - Encoding a Geo-Holdout as a Bayesian Experimental Design and Computing Its EIG
- Identity Fragmentation and the Privacy-Era Limits of User-Level Tests — contamination across a user’s devices and identities as a form of interference