Identity Fragmentation and the Privacy-Era Limits of User-Level Tests

Summary

A user-level experiment presumes that the unit randomized, the unit exposed and the unit whose outcome is measured are the same person. On the open web they are not: cookies are “browser, device, and site specific”, device IDs miss cross-device journeys, and privacy changes (third-party-cookie blocking, IDFA opt-in) are removing the linking tools. Lin & Misra (2022) characterise the resulting identity fragmentation bias in a linear model as the sum of three terms — purchase fragmentation (always attenuating), exposure fragmentation (an omitted-variable term of either sign) and spurious covariance (device-level activity bias and cross-device substitution, either sign). Contrary to the folk wisdom that fragmentation merely attenuates, the total “cannot be signed or bounded under standard assumptions. Instead, upward biases and sign reversals can occur even in experimental settings.” Only under symmetric and independent exposures — deliverable by a fully randomized ITT design — does the bias reduce to division by the number of fragments. Partial identity linking can make things worse; aggregation to groups that contain all of a person’s fragments — in the limit, geographies — is unbiased at the cost of power.

Overview

Earlier notes in this cluster treat identity as solved. Johnson, Lewis & Nubbemeyer acknowledge it is not: in their cookie-based Google Display Network test, cookie churn loses post-exposure purchases, a cookie-switching consumer “could switch treatment groups”, and “each computer, browser, tablet, or mobile device has its own unique anonymous cookie that is independently randomized. Hence, if a user’s ad exposures are linked to one cookie but her purchases are linked to another, our estimates will be further attenuated” — so they “interpret the absolute ad lift estimates as lower bounds” (§5.3). Gordon et al. (2019, fn. 3) list the two consequences of cookie-level data: “users in an experimental control group may inadvertently be simultaneously assigned to the treatment group”, and “advertising exposure across devices may not be fully captured”; Facebook avoids both only because users must log in everywhere.

Lin & Misra formalise what is lost, and show the “lower bound” reading is safe only in special cases.

Scale of the problem (their §1): Facebook reported 32% of users who show interest in mobile ads convert on desktop; cookie syncing “can match only up to 60% of fragmented records”; iOS 14 made IDFA collection opt-in; Safari and Firefox block third-party cookies and Chrome had announced it would follow. (Chrome’s deprecation timeline subsequently slipped and changed form, but the direction the authors describe — “identity fragmentation is here to stay” — has held.)

Main Content

Framework

Fragmented data (Lin & Misra §3.1) ^def-fragmented

True model at the person level, with two devices :

The analyst sees “users”: stacked outcomes and exposures . A person buys on device 1 with indicator (on device 2 otherwise), so with , and is the device-preference matrix. Assumptions: and standard OLS conditions, so that any remaining bias is due to fragmentation alone.

Bias decomposition (Lin & Misra eq. 8) ^thm-bias-decomposition

With the (positive-definite) lower-right block of ,

  • — purchase fragmentation. Each row captures only part of the outcome variation while the row count doubles: pure attenuation, . Always present unless .
  • — exposure fragmentation. Exposure on the other device is an omitted variable (Omitted Variables Bias); sign and size follow the cross-device exposure correlation . Vanishes when exposures are independent across devices and mean-centred.
  • — spurious covariance. Proportional to the baseline and to , the covariance between where ads are seen and where purchases happen. Positive under device-level activity bias (people see ads and buy on the same device); negative if ads are seen on the phone and purchases made on the desktop; also generated by cross-device substitution (an ad shifts a purchase between devices without changing the total).

“The latter two bias components have arbitrary signs and magnitudes … Moreover, this bias does not converge to zero in the limit.” is bounded in , so only when the baseline (e.g. new products) is with the correct sign. “In most digital ad effect studies, ” — the low signal-to-noise regime of The Unfavorable Economics of Ad Experiments - Power and Signal-to-Noise — and dominates.

Device-level activity bias is removable by estimating separate models per device (“unstacking”); cross-device substitution is not (§3.2). The results extend to device-specific coefficients, fragments, and mixtures of fragmented and complete identities (§3.4-3.5), and the sources of bias persist in GLMs such as logit and Poisson, where attenuation is if anything stronger (§6).

Does randomization rescue the estimate?

Symmetric and independent exposures (SIE) (Lin & Misra §3.3, §4.2) ^thm-sie

Suppose (1) , and , and (2) does not depend on exposures. Then with mean-centred covariates ; with fragments per user , and is unbiased, with a confidence interval “at least times” that of unfragmented data. Coey & Bailey’s (2016) people-and-cookies correction is the special case of i.i.d. cookie exposures and equiprobable purchase cookies.

Both conditions are needed. With randomized () but asymmetric treatment intensity — , , , , , constant device preference — the bias is :

“Randomization by itself does not provide additional guarantees regarding the sign of the fragmentation bias.”

Implications for experiment design:

  • Randomize at the fragment level with equal probabilities across fragment types and analyse as ITT. Condition (1) then holds by design, and condition (2) holds for assignment even though it fails for realized exposure (activity bias). “We recommend researchers focus on intent-to-treat.” The ATT scaling of From ITT to Treatment-on-the-Treated in Ad Experiments is then contaminated twice: through the ITT numerator (by under SIE) and through an exposure rate measured on fragments rather than people.
  • Contamination is attenuation plus. If one person’s cookies fall in different arms, “control” people are partly treated. Under SIE this is the factor; without symmetry (most mobile inventory in one arm’s reach, most purchases on desktop) the direction is not guaranteed.
  • varies across people, and the corrected estimator “will assign more weights to users with more fragments … likely users with more activities.”

Three remedies compared (Lin & Misra §4)

RemedyIdeaStrengthWeakness
Identity linkingDeterministic (login, hashed email) or probabilistic (fingerprint, cookie sync) stitchingIntuitive; industry defaultOnly partial in practice. The pooled estimator is with matrix weight ; bias need not fall monotonically in the linked share, and in simulations intermediate linkage is “often further away from , and sometimes the bias takes the reversed sign”. Report match rates.
Experiment-based adjustmentMultiply the ITT by under SIESimpleNeeds full randomization across fragment types, a common-effect model, known ; wider intervals
Stratified aggregationAggregate to groups guaranteed to contain all fragments of each person — geography (county, city, store), refined by covariate cells such as zipcode × gender × age × education”Requires the least assumptions”; works for nonlinear and structural models; no linking or experiment neededLess power, less heterogeneity; sensitive to error in the binning covariates

The third row is the formal bridge to geo experiments: “A simple form of aggregation is analyzing data at the geographic level … While simple aggregation provides robustness, it significantly reduces statistical power.” A geo test is the design in which every fragment of a person shares one treatment assignment (they live in one place) and outcomes are summed over all fragments and channels, so are zero by construction.

Empirical illustration (Lin & Misra §5)

Matched person-level data from an online durable-goods seller (391,195 consumers; display, search and social exposures; engagement outcome), artificially fragmented into mobile / desktop / tablet records. Ads and engagements concentrate on the same devices; own-device ad-outcome correlations are positive (0.06-0.24) while cross-device correlations are negative — the signature of device substitution.

CoefficientTrue (person level)FragmentedRatio
Search0.25840.36261.40
Social0.14390.21571.50
Display0.04280.08061.88

Fragmentation inflates effects by 40-88% — the opposite of the attenuation intuition — and, because the row count triples, standard errors shrink, leaving “zero overlap between the fragmented and true estimates”. Stratified aggregation on simulated MSA × age × income cells gives intervals that are wider but cover the truth (Fig. 3).

Identity tiers in practice

Barajas, Bhamidipati & Shanahan’s tutorial outline ranks the units available to an incrementality platform:

  1. Cookie-based — the DSP display standard; “cookie deletion, one per device”.
  2. Device-ID based — standard for in-app and open-exchange buying; “id reset, normalization, hashing”.
  3. Logged-in users — “common within the walled gardens”; cross-device execution and attribution; “as reliable as product experimentation”.
  4. Household-level — identity graphs built on IP clustering; “IP churn and reset, traveling, moving”.

The privacy era removes tiers 1-2 for a growing share of traffic, leaves tier 3 to platforms with login walls — Lin & Misra’s concluding worry about “a stronger incumbent advantage”, echoing Lewis & Rao’s scale argument — and makes tier 4 and geography the fallbacks for everyone else.

Examples

Lin & Misra’s Table 1, reproduced. Two people, each with a desktop (D) and mobile (M); both buy one unit; person 1 saw 2 ads, person 2 saw 4. True .

import numpy as np
slope = lambda x, y: np.polyfit(x, y, 1)[0]
 
# person level: ads (2, 4), purchases (1, 1)  -> slope 0
print(slope([2, 4], [1, 1]))                        # 0.0
 
# (b) ads seen mostly on the purchase device
print(slope([2, 0, 3, 1], [1, 0, 1, 0]))            # +0.4
 
# (c) ads seen on mobile, purchases on desktop
print(slope([0, 2, 1, 3], [1, 0, 1, 0]))            # -0.4

The same zero effect reads as strongly positive or strongly negative depending only on the cross-device pattern of exposure and purchase — in action. Aggregating rows back to the person (or to any group containing both devices) restores the zero.

A quick diagnostic from their §5.2. Without cross-device substitution, the matrix of correlations between device-level outcomes and device-level exposures has proportional columns. Large positive diagonals with negative off-diagonals, as in their Table 2, flag substitution and hence upward bias in cookie- or device-level attribution.

Connections

See Also