From ITT to Treatment-on-the-Treated in Ad Experiments

Summary

In an ad holdout, assignment is random and exposure is not. Because held-out users are never exposed, noncompliance is one-sided: the population splits into compliers (would be exposed if eligible) and never-takers, with no always-takers or defiers. Under an exclusion restriction — assignment affects outcomes only through exposure, plausible because users do not know their arm — the effect on exposed users is the intent-to-treat effect divided by the exposure rate, . This is the Wald / 2SLS estimator with as the instrument, and here the LATE is the ATT. Lift re-expresses the ATT relative to the exposed users’ counterfactual conversion rate. The scaling changes the point estimate, not the -statistic: precision gains come only from designs that identify the control arm’s would-be-exposed users.

Overview

Gordon, Zettelmeyer, Bhargava & Chapsky (2019, §3) give the cleanest statement of the estimands, in the Imbens-Rubin potential-outcomes notation also used in Potential Outcomes Framework. Each study has users , random assignment , potential exposure , covariates , and potential outcomes (converted or not). The defining feature of a holdout is

“compliance is perfect for users in the control group, who are never shown campaign ads. However, compliance is one-sided in the test group, where exposure (receipt of treatment) is an endogenous outcome that depends on factors related to the user, platform, and advertisers” (§2.3). Three user groups are observed: control-unexposed, test-unexposed, test-exposed.

Why exposure is endogenous even inside a randomized test arm (§2.3):

  • User-induced — activity bias: you must be online to be exposed, and being online predicts online conversion.
  • Targeting-induced — the delivery system up-weights bids for users it predicts will convert; optimizing for clicks can even produce negative selection on purchases (“clicky users”).
  • Competition-induced — winning an auction depends on rivals’ bids; rivals selling complements who win impressions can leave the unexposed pool enriched with likely buyers.

Randomization neutralizes all three for the contrast because the same bid-weighting is applied to both arms; “for members of the control group, the focal ad is replaced ‘at the last moment’ by the runner up.”

Main Content

Assumptions ^def-assumptions

  1. SUTVA — one version of treatment; no interference between users. Supported on Facebook by single-user login (no accidental exposure of controls) and by users not knowing their arm; ad-sharing from test to control users would make estimates conservative.
  2. Random assignment — . Supports an ITT analysis on its own.
  3. Exclusion restriction — for : “assignment affects a user’s outcome only through receipt of the treatment. Because users are unaware of their assignment status, only exposure should affect outcomes.” Required only for the ATT.

ITT and ATT ^def-itt-att

The ITT “should be interpreted as conditional on the platform’s ad-optimization system” (fn. 10) — the “entire treatment” includes who the delivery engine chose to reach. The ATT “is inherently conditional on the set of users who end up being exposed”, so it is not comparable across campaigns with different targeting.

ITT-to-ATT scaling under one-sided noncompliance (Gordon et al. 2019 eq. 5-7) ^thm-itt-att

Under assumptions 1-3,

where is the share of compliers, since .

Derivation. Write the ITT as a mixture over compliers and never-takers,

For never-takers , so by exclusion . Hence . “Scaling by the inverse of ‘undilutes’ the ITT effect.” Imbens & Angrist call this quantity the LATE; “if the sample contains no ‘always-takers’ and no ‘defiers,’ which is true in our experimental design with one-sided non-compliance, the LATE is equal to the ATT.”

In practice the ATT and its standard error come from two-stage least squares of on with as the instrument (Gordon, Moakler & Zettelmeyer §4.1). Johnson, Lewis & Reiley call the same thing the indirect TOT estimator and note it “is numerically equivalent to computing a local average treatment effect by using the random assignment as an instrument for treatment” (fn. 2). The first stage is as strong as an instrument gets — is typically 0.3-0.8 and estimated from millions of users — so weak-instrument concerns do not arise; the difficulty is entirely the noisy reduced form.

Lift ^def-lift

the incremental conversion rate among treated users as a percentage of “the estimated conversion rate of the treated group if they had not actually been treated” (Gordon et al. 2019 eq. 8). The denominator is not the control mean: exposed users are selected, so their counterfactual baseline differs from the arm-wide control rate. Lift normalizes across advertisers and outcomes but “differences between methods can seem large when the treated group’s baseline conversion rate is small”. Confidence intervals are obtained by bootstrap because lift is a ratio.

What the scaling does and does not buy

Scaling does not improve precision; identifying the counterfactual treated does ^thm-precision

Since is estimated almost without error, and : the -statistic is unchanged. With users per arm and common outcome variance ,

where the direct estimator compares treated users with identified counterfactual-treated controls (placebo- or ghost-tagged). The variance ratio is . For this predicts a 25.6% smaller standard error, matching the 25% Johnson, Lewis & Reiley measure; for the 3% exposure rates of on-the-fly ITT tests it is a 33-fold variance penalty. Johnson, Lewis & Nubbemeyer’s empirical ITT-to-PGA variance ratios of 5.9-17.1 additionally reflect pruning of pre-exposure outcomes.

So there are two distinct estimators of the same ATT:

Indirect (Wald / 2SLS)Direct (control-ad or ghost-ad tagged)
NeedsAssignment, exposure in test armA symmetric exposure flag in both arms
SampleEveryone eligibleFlagged users only; post-first-exposure outcomes
Bias riskExclusion restriction onlyAsymmetric tagging (PSA under optimization; PGA under-prediction)
PrecisionBaselineVariance reduced by about , more with outcome pruning
Used byMeta Conversion Lift; geo testsYahoo! PSA tests; Google predicted ghost ads

Interpretation cautions

  • The ATT is local to the delivery system. Compliers are whoever the auction, pacing and targeting model chose to reach at that budget. “If campaign length and budget were increased, additional unexposed users might become exposed” (Gordon, Moakler & Zettelmeyer fn. 8). Extrapolating an ATT to a larger budget assumes marginal compliers respond like average ones — usually optimistic, since delivery systems reach the most responsive users first.
  • It is conditional on everything else running, “such as marketing activities the advertiser conducts in other channels (e.g., search, TV) and its competitors’ activities” (2019 §2.2).
  • Short windows give conservative totals. Conversions after the measurement window are missed (2019 fn. 19); see Delayed Feedback Model for Conversion Prediction for the censoring structure, and The Unfavorable Economics of Ad Experiments - Power and Signal-to-Noise for why long windows hurt power.
  • “Non-compliers” did not choose. Unlike a drug trial, unexposed users “did not fail to comply as a result of their own deliberate decisions”; exposure results from user activity, the auction and the platform jointly (2023 §1). The never-taker/complier language is formal, not behavioural.
  • ITT or ATT for decisions? Total incremental conversions are identical either way: . ITT answers “what did launching this campaign at this audience buy?”; ATT answers “what is an exposure worth?” and is the quantity observational methods try to estimate, which is why Gordon et al. benchmark on it (Experimental Benchmarks for Observational Ad Measurement).

Examples

Study 4 of Gordon et al. (2019), a retail checkout outcome (Tables 3-4, fn. 20), 25.6M users, 70/30 split:

QuantityValue
Control conversion rate0.033%
Test conversion rate0.045%
ITT0.012 pp (ITT lift 37.7%, CI [27.2%, 49.1%])
Share of test users exposed, 37%
ATT 0.033 pp
Exposed-in-test conversion rate0.079%
Counterfactual rate of exposed 0.046%
ATT lift72.8%, CI [49%, 103%]
Unexposed-in-test conversion rate0.025%

The counterfactual rate for exposed users (0.046%) follows from the identifying assumption that unexposed test users would have converted at the same rate in control: gives . Exposed users’ baseline is nearly double the unexposed users’ — the selection that makes the naive exposed/unexposed comparison report a lift of 316%.

import numpy as np   # y, z, w are user-level numpy arrays
 
def lift_from_holdout(y, z, w):
    itt = y[z == 1].mean() - y[z == 0].mean()
    pi_co = w[z == 1].mean()                 # w is 0 for all z == 0 by design
    att = itt / pi_co                        # == 2SLS coefficient on w, instrument z
    base_exposed = y[(z == 1) & (w == 1)].mean() - att
    return dict(itt=itt, pi_co=pi_co, att=att, lift=att / base_exposed)
 
# bootstrap users (not conversions) for the lift interval

Connections

See Also