Activity Bias: Why Observational Ad Measurement Flatters Itself

Compare the people who saw your ad with the people who didn't, and advertising looks miraculous. Lewis, Rao, and Reiley ran that comparison against randomized ground truth and found the observational estimate overstated the true effect by two orders of magnitude — and adding every control they had barely moved it. This is the founding cautionary tale of modern ad measurement, and it is worth knowing in detail.

This post opens a series on the research that shapes how we think about measurement. It starts here because the failure it documents is the one every other method in the series exists to avoid: reading an exposure correlation as a causal effect. The paper's numbers are extreme enough that once you have seen them, you stop asking whether observational lift studies are biased and start asking only how badly.

The Comparison Everyone Wants to Make

Digital advertising generates an irresistible dataset. The platform logs who was served each impression; the analytics stack logs who searched, browsed, signed up, and bought. Join the two, split users into exposed and unexposed, and difference the outcomes. The analysis is cheap, it is always available, and at the scale of a major campaign it is precise — millions of users make the confidence intervals gratifyingly narrow.

The trouble is the thing the join quietly assumes: that seeing the ad was, conditional on whatever covariates you include, as good as random. It was not. To be exposed to a display campaign you had to be online that day, on the specific pages where the campaign ran, during the window when it flew. The population that satisfies those conditions is not a random draw of your customers — it is the population that happened to be active online at that moment. And people who are active online do more of everything online.

The Mechanism: Activity Bias

Definition: Activity bias

The systematic overestimation of advertising effects that arises because ad exposure is itself a marker of online activity. Users who see a campaign are, by construction, active on the days and pages where it runs — and those same active users search more, view more pages, sign up for more services, and buy more, ad or no ad. An exposed-vs-unexposed comparison therefore attributes correlated background activity to the advertising (Lewis, Rao, & Reiley, 2011).

What makes activity bias so destructive is that the confounder is not a stable trait you could measure once and adjust away. It is a day-level state: being an active internet user today. Ad exposure on a given day is close to a proxy for that state, which means exposure and the confounder are almost the same variable. Conditioning on demographics, historical behavior, or even same-day usage summaries leaves most of the entanglement intact, because none of those observables fully captures the fluctuating propensity to be online and doing things at the moment the campaign is in flight.

In the language of identification, the conditional independence assumption fails, and fails badly: there is no set of observed covariates rich enough to make exposure ignorable. That is a structural problem, not a data-quality problem, which is why the fixes discussed later are design changes rather than better regressions.

Three Experiments, One Lesson

The paper's contribution is not the abstract argument — it is that the authors could check the observational answer against experimental truth. Working with large randomized campaigns on Yahoo!, they measured the same effects both ways: once using the randomization, and once pretending they only had observational data, exactly as an analyst without an experiment would.

Experiment 1: Brand searches

The first study measured whether a large display campaign on the Yahoo! front page lifted brand-related searches. The randomized comparison — treatment versus a control group that was withheld the ad by design — put the true effect at a 5.4% increase in brand searches. Real, positive, and modest.

The observational version of the same question compared users who happened to be exposed with users who were not. With no controls, the estimated lift was 1,198% — the exposed group searched for the brand at roughly thirteen times the rate of the unexposed group. That is not a subtle discrepancy; it is an estimate about 220 times larger than the truth. And, as the next section shows, throwing controls at it recovered almost none of the gap.

Experiment 2: Page views

The second study tracked page views on the advertiser's own Yahoo! content around a similar campaign. The pattern repeated: observational comparisons showed large apparent effects that the experimental benchmark did not support. Notably, an apparent “competitive effect” — activity that looked like it was being driven to related content by the ad exposure — turned out to be minimal once the randomization spoke. Exposed users visited more of everything, and a naive analysis read that as the campaign radiating influence across the site.

Experiment 3: Sign-ups at a competitor

The third analysis is the most instructive, because it works like a placebo test. The outcome was account sign-ups at a competitor's website — something the campaign had no plausible business driving. Exposed users nevertheless showed a clear spike in competitor sign-ups on the campaign day. The tell: the randomized control group, which never saw the ad, showed the same spike on the same day. The entire correlation was ambient activity. People were simply more active online that day — visiting Yahoo!, seeing ads, and signing up for services everywhere — and the exposed-vs-unexposed split repackaged that co-movement as an advertising effect.

Three different outcomes — searches, page views, sign-ups — and one mechanism underneath all of them. Activity bias is not specific to a metric; it contaminates every outcome that active users produce more of, which is to say nearly every outcome an advertiser cares about.

Why Controls Don't Fix It

The instinctive response to a confounded comparison is to add controls. The paper runs that playbook to its end, and the progression is the most quotable table in the ad-measurement literature. For the brand-search outcome, where the experimental truth was 5.4%:

ModelControlsEstimated lift
(0)None1,198%
(1)Day dummies894%
(2)+ Session dummies871%
(3)+ Page views, minutes spent872%
TruthRandomized experiment5.4%

Each row adds information an analyst would reasonably reach for: which day it was, which session, how many pages the user viewed, how many minutes they spent. The estimate falls from 1,198% to the high 800s and then simply stops falling — the last, richest specification lands at 872%, still roughly 160 times the experimental answer. The controls absorb a sliver of the activity signal and leave the rest fused to exposure.

Matching fares no better, and for the same structural reason. Propensity score matching promises to pair each exposed user with an observably similar unexposed one. But in display advertising, exposure is determined by visiting particular pages at particular times — so the “similar” unexposed user is someone who looked alike on the covariates yet, in that moment, was doing something different online. The matched groups differ in precisely the unobserved dimension that drives the outcome. This is the selection problem in its purest form: units select into treatment on the very trait the treatment is supposed to move. The assumptions that would license a causal reading here belong to the same family this site catalogs for aggregate models on the identification assumptions page — and in this setting they are simply false.

The wider the activity overlap, the wilder the estimate

The randomized truth is fixed: advertising lifted brand searches by 5.4%. The observational estimate compares exposed users to unexposed ones — but exposure is a marker of being active online, and active users search more anyway. Drag the activity–exposure correlation. As exposure and “active today” become nearly the same variable, the observational lift explodes toward the paper’s 1,198% while the true incremental line stays flat along the floor.

0.85
Observational lift
True incremental lift5.4%
Inflation factor
Reading

Observed lift = 5.4% + (1,198% − 5.4%) · r2.2. At r ≈ 0.85 the estimate lands near 840% — roughly the same order as Lewis–Rao–Reiley’s fully-controlled 872%, still ~155× the experimental truth. The gap is bias, not variance: more users only sharpen the wrong number.

⚠️ More data makes it worse, not better

Activity bias is a bias problem, not a variance problem. Scaling from thousands of users to millions shrinks the standard errors around the wrong number, turning a flattering estimate into a flattering estimate with four significant digits. No sample size rescues an identification strategy that conflates exposure with activity — precision is not validity.

💡 A textbook Type M error

In the error taxonomy of Gelman and Carlin (2014), the 1,198%-vs-5.4% comparison is an exaggeration ratio of roughly 220× — a magnitude (Type M) error, delivered with high statistical confidence. It is a useful mental benchmark: when someone quotes an observational ad lift, the honest uncertainty is not the confidence interval on the estimate but the unmodeled gap between the estimate's assumptions and reality, and that gap has been measured at two orders of magnitude.

What This Means for Practice

The practical conclusions follow directly from the mechanism, and they have held up in the fifteen years since the paper circulated.

Randomize when you can. The experiments in the paper did not merely produce better estimates — they produced the only credible estimates, because randomly withholding the ad from a control group is the one operation that severs the link between exposure and activity. The control group spikes on the campaign day too, and the design subtracts that spike out. Everything this site builds on experimentation — geo holdouts, calibration experiments, sequential testing — traces its justification to this point.

When you cannot randomize, name your identification strategy. Not every effect can be tested, and aggregate methods like media mix modeling exist precisely because user-level randomization is often unavailable. But the lesson transfers: an observational method earns a causal reading only through explicit, stated, defensible assumptions — a strategy for why the variation you are exploiting is as good as random — never through the volume of controls in the regression. The causal inference page lays out what that looks like for MMM: structural assumptions, confounders named up front, and experimental calibration wherever the assumptions are weakest. The distance between that posture and “we added controls” is exactly the distance between 5.4% and 872%.

Treat exposed-vs-unexposed lift studies as upper bounds at best. A vendor-reported lift built on comparing exposed users to matched unexposed users inherits activity bias in full. Such numbers should not be used to benchmark experimental results or MMM outputs — when they disagree with a randomized readout, the disagreement is the expected direction and the expected sign, not evidence that the experiment “missed” something.

Expect the mechanism outside advertising. Activity bias is a special case of a general pattern: whenever exposure to a treatment correlates with baseline engagement — email opens, push notifications, feature announcements, in-app promotions — the exposed population is selected on the very activity that produces the outcome. The 100×-plus magnitudes documented here are what that selection can do when the confounder and the exposure are nearly the same variable.

None of this says advertising doesn't work. The randomized estimate in Experiment 1 was positive — advertising moved brand searches by a real 5.4%. What the paper demolishes is a measurement practice, not a channel: the observational comparison didn't exaggerate a zero, it exaggerated a genuine effect by a factor of a couple hundred, which is arguably worse, because the flattering number is always available and always larger than the honest one.

Takeaways

  • Exposed-vs-unexposed comparisons conflate ad effects with online activity: the same users who see ads search, browse, and sign up more everywhere, ad or no ad.
  • The measured magnitude is enormous — a 1,198% observational lift against a 5.4% experimental truth, with the fully controlled model still ~160× too high.
  • Controls and propensity matching cannot repair the comparison, because day-level activity is entangled with exposure in ways no observable covariate set captures.
  • More data sharpens the wrong answer: activity bias is a bias problem, and sample size only buys precision around a systematically inflated estimate.
  • Credible ad measurement requires randomized experiments or an explicit, defensible identification strategy — never a raw exposure correlation.

References