The Platform Randomizes After You Do
Two creatives, split-tested inside a platform's own experimentation tool. Users randomized 50/50 by the platform, assignment logged, no self-selection into cells, no analyst degrees of freedom in the readout. By every standard this series has been applying, that is a randomized controlled trial. It still produces a confounded comparison. The reason is that your randomization is not the last randomization to happen. After users are split, the delivery algorithm decides which of them actually sees each creative, and it makes that decision separately per creative, optimizing each one against its own predicted-response model. The two cells end the test having reached measurably different audiences. The measured contrast is the ad content plus the algorithm's selection, and nothing in the log tells you the split. Braun and Schwartz established this for ad A/B tests in 2025. Also in 2025, a Meta audit of 3,204 lift tests against 181,890 A/B tests measured it at scale. The corollary neither paper leads with is the operationally useful part, and it says something much narrower than "platform experiments are broken." In a lift test both cells face the identical delivery algorithm, so the algorithm cancels and the test stays valid. In a content A/B the algorithm is part of the treatment. So the sharp version of the claim is a registry rule: a creative A/B result is inadmissible as an MMM calibration input, and because platform tests are enormous, it will arrive with a standard error small enough to overrule the geo holdout that was valid.
Randomization Was Not the Problem
This series has spent several posts on the failure mode where the exposed group differs from the unexposed group because users chose exposure. Activity bias is the canonical version: people who saw an ad were online, and being online predicts converting, so the exposed-versus-unexposed contrast overstates the ad by orders of magnitude and no set of controls repairs it. Attribution is not incrementality makes the same point through the bookkeeping: last-touch credit conditions on the outcome path, and the fix in both posts is the same word. Randomize.
That advice was right and this post does not retract it. What follows is a different mechanism that happens to survive the fix. Here the platform does the selecting, deciding per treatment arm which of your randomized users get treated at all. The assignment is clean. The delivery is not, and the delivery is where audience composition gets decided.
The structural parallel is with a sibling post in this series rather than with the observational ones. Randomization buys you the total effect; the funnel is not included argues that a randomized geo lift identifies one contrast and is silent about another one printed beside it. This post is the same shape of error at a different layer: randomizing users identifies the effect of the campaign as delivered and is silent about the effect of the content, which is the quantity the creative team believes it just measured.
What the Algorithm Does After You Randomize
Start with what a platform delivery system is for. Every dollar it places is placed on purpose, never by coin flip. It is a ranking system that predicts, for every eligible user and every candidate ad, the probability that this user responds to this ad, and then spends the budget on the highest predictions. That is the product working correctly. It is also, during an experiment, a second assignment mechanism operating downstream of yours and optimized to a different objective.
Definition: divergent delivery
Braun and Schwartz (2025) name the phenomenon divergent delivery: when an experimenter supplies two or more ads and lets the platform optimize each of them, the platform delivers each ad to a different subset of the eligible users, chosen to suit that ad. Exposure within the test is therefore not random, even though assignment was, and the estimated comparison confounds the effect of the ad content with the effect of algorithmic targeting. The paper's own emphasis is that this is not a bug the platform will patch. Undetectable optimization is the service being sold. The mixes of users differ in ways the experimenter cannot observe from the reported data.
The corresponding empirical literature is older than the term. Ali and colleagues (2019) showed that on Facebook, holding the advertiser's targeting parameters fixed and varying only the creative and the budget produced large, systematic skews in who was reached: along gender and racial lines, for real employment and housing ads, with neutral targeting. The delivery stage, not the targeting stage, generated the skew.
To see what that does to a content comparison, take the cleanest possible case and grant every assumption the experimenter wants. One dimension of user heterogeneity, call it responsiveness \( x \), standardized so \( x \sim N(0,1) \) in the eligible pool. A user's baseline propensity to convert rises with it. Two creatives whose content effects are known and which differ by a real amount \( \delta \). Creative B is the better ad, by construction. Randomize the pool 50/50. Now let the delivery system spend each cell's budget on the users it scores highest for that cell's creative, with the two creatives appealing to opposite ends of \( x \): a discount execution and a premium execution, say, which is exactly the case where creative testing is most interesting and where the algorithm has the most reason to diverge.
The measured contrast is then the true content difference multiplied by a factor that depends only on how hard the algorithm optimized. And the natural way to see the damage is in the units the auditors use: the standardized mean difference between the two cells' reached audiences.
Divergent delivery, with the truth pinned
Creative B is the better content by the amount on the second slider. That is a fact of the simulation, drawn as the flat dashed line. The horizontal axis is the resulting audience imbalance between the two cells, measured as a standardized mean difference on the responsiveness dimension. Both cells reach 250,000 users at a baseline conversion of 1.0%. Push the optimization strength up and watch the measured contrast leave the truth, cross zero, and declare the worse creative the winner, with an interval that never widens.
The curve is exact, a closed form rather than a Monte Carlo draw (see the deep dive). At the defaults the truly better creative is measured at −9.8% instead of +8%, and the sign flips at an imbalance of SMD = 0.13. That sits below the 0.20 bar the Meta audit treats as the threshold for meaningful imbalance. The band is the 95% sampling interval. It is narrow everywhere, which is the whole problem.
Deep diveWhy the bias is a population quantity
Let responsiveness be \( x \sim N(0,1) \) in the eligible pool and let a user's baseline conversion propensity be \( b(x) = b_0 \exp(\gamma x - \gamma^2/2) \), normalized so that \( \mathbb{E}[b] = b_0 \). Cell \( k \) reaches user \( i \) with probability proportional to \( \exp(\lambda\theta_k x_i) \), where \( \lambda \ge 0 \) is how aggressively delivery optimizes and \( \theta_A = +1 \), \( \theta_B = -1 \) encode that the two creatives are scored toward opposite ends of the dimension.
Exponentially tilting a standard normal shifts it and nothing else: \( \phi(x)e^{\lambda\theta x} \propto \phi(x - \lambda\theta) \). So the audience actually reached in cell \( k \) is \( x \sim N(\lambda\theta_k,\, 1) \), same variance, shifted mean. Using \( \mathbb{E}[e^{\gamma x}] = e^{\gamma\mu + \gamma^2/2} \) for \( x \sim N(\mu, 1) \),
$$ \mathbb{E}\!\left[\,b \mid \text{reached in } k\,\right] \;=\; b_0\, e^{\gamma\lambda\theta_k}. $$With a true multiplicative content advantage \( \delta \) for B, the observed rate ratio is \( (1+\delta)\,e^{-2\gamma\lambda} \). The standardized mean difference of \( x \) between the two reached audiences is \( d = 2\lambda \), because both have unit variance. Substituting,
$$ \widehat{\text{contrast}} \;=\; (1+\delta)\,e^{-\gamma d} - 1, \qquad \text{sign flips when } \; d \;\gt\; \frac{\ln(1+\delta)}{\gamma}. $$Every quantity on the right is a population expectation. There is no \( n \). As the test grows, the estimator converges to the content effect times \( e^{-\gamma d} \), and never to the content effect itself. That is inconsistency, not imprecision, and it is why the next figure looks the way it does. At \( \delta = 0.08 \) and \( \gamma = 0.6 \) the flip threshold is \( \ln(1.08)/0.6 = 0.128 \).
One honest caveat on the units. Burtch and colleagues compute their standardized mean differences on observable user features, whereas the \( x \) above is the latent responsiveness dimension itself, which is the worst case. An imbalance of 0.13 on an observed feature only loosely related to response does proportionally less damage. Cutting the other way, and this is Braun and Schwartz's point: the dimension the platform tilts on is by construction the one it predicts response from, and it need not be visible in the features you are able to audit.
Why the Lift Test Survives
Everything above is an argument about comparing two ads. Notice what it does not touch. In a lift test the comparison is between users eligible to be shown the campaign and a held-out control group shown nothing. There is one ad configuration, and one delivery model optimized to one objective. The algorithm still selects which of the eligible treatment users get reached. But the control group's counterfactual counterparts are defined by that same selection, either by holding out at randomization and comparing intent-to-treat, or by logging the ad a control user would have won. The algorithm applies identically on both sides of the contrast, so it cancels. Nothing about divergent delivery threatens a holdout.
This has been measured. Burtch, Moakler, Gordon, Zhang, and Hill (2025) audited Meta's own experimentation surface across 3,204 lift tests and 181,890 A/B tests, computing standardized mean differences in user characteristics across cells and adopting Cohen's conventional 0.20 threshold for meaningful imbalance. Their summary is worth quoting exactly, because the two halves are usually reported separately and the pairing is the finding: "Lift tests show no meaningful audience imbalance, confirming their causal validity, while A/B tests show clear imbalance, as expected." The magnitudes are stark. Among lift tests, 0.16% of standardized mean differences exceeded 0.20 and 5% of the balance t-tests were significant at the 5% level. Exactly the nominal rate. That is what randomization predicts, and what a clean test looks like. Among A/B tests, 22% of standardized mean differences exceeded 0.20 and 25% of the t-statistics were significant: five times the nominal rate.
💡 "As expected" is doing real work in that sentence
The authors are emphatic that divergent delivery in A/B tests is intentional. In their words, it is "specific to A/B tests and intentional, informing advertisers about ad performance in practice." The imbalance is the algorithm doing its job. That framing matters for how you write the caveat: an A/B test is a valid measurement of a different estimand. It tells you which configuration performed better as the platform will actually deploy it. Operationally that is useful. It is also a different quantity from the content effect. Braun and Schwartz reach the same prescriptive place from the other direction: use these tests to predict performance within the same platform and campaign, and stop short of causal claims about the creative that you intend to carry anywhere else.
So the asymmetry is not a hedge. Randomized platform holdouts remain the strongest causal anchor an advertiser can buy, and that endorsement, made in an earlier post, stands unqualified. The A/B is the special case that fails, and it fails for a reason specific to having two treated cells and no control.
Sample Size Is Not a Remedy
The reflex response to any measurement problem on a platform is scale. Platforms have hundreds of millions of users, so run the test longer, reach more people, and the noise goes away. It does go away. That is the problem, because the noise was never what was wrong.
Write the measured contrast as truth plus bias. The bias term derived above, \( (1+\delta)e^{-\gamma d} - (1+\delta) \), contains no sample size. The standard error falls like \( 1/\sqrt{n} \). Their ratio therefore diverges, and the test passes through three regimes as it grows. First the interval covers the truth and also covers everything else. Then it stops covering the truth. Then it excludes zero on the wrong side, and the platform reports a statistically significant winner that is the worse ad. Scale does not move you toward the answer. It moves you from uninformative to confidently wrong, and it does so at sample sizes that are small by platform standards.
The futility panel: bias is flat, precision is not
Same simulation, swept over the number of users reached per cell. The flat line is the absolute bias in percentage points, and the falling one is the half-width of the 95% interval. Where they cross, the interval stops covering the true contrast. True content advantage fixed at +8%, responsiveness spread γ = 0.60.
At the defaults the bias is −17.8 pp and never moves. About 19,000 reached users per cell is enough for the 95% interval to exclude the true contrast, and 62,000 is enough to report the wrong creative as a significant winner. Platform tests routinely run two to three orders of magnitude larger than that.
This is the same arithmetic that makes an underpowered experiment and a biased experiment fail in opposite directions, and it is why "we ran it on ten million users" is not a defense. It is also why the failure is socially durable: the biased estimate arrives with a tight interval and a small p-value, which is precisely the presentation that survives a review meeting. A p-value cannot tell you which of two numbers is closer to the truth, and here it reliably certifies the wrong one.
Why It Cannot Be a Calibration Input
Now the part that reaches an MMM. The experiment is the prior argued that a randomized lift test is the right way to anchor a weakly identified channel, because Bayesian updating is additive in precision: the observational likelihood carries little information in a collinear channel's direction, so a reasonably tight experimental measurement dominates and drags the channel to a value someone actually measured. That argument is correct, and it has an edge nobody likes to look at. Precision-weighted pooling has no bias term. It weights inputs by \( 1/\text{se}^2 \) and by nothing else. Hand it a biased measurement with a small standard error and it will do exactly what it promised to do.
The path from a creative A/B to an MMM calibration slot is short and entirely plausible. A team split-tests two Social creatives, the platform declares a winner, and the winning cell's reported return (3.1×, on tens of millions of impressions, with a standard error in the second decimal place) gets written into the experiment registry as a measured readout for the Social channel, because the test was randomized and the number did come from an experiment. Meanwhile the geo holdout that ran last quarter measured the same channel at 1.6× with a standard error of 0.35, because geo holdouts have a few dozen markets and a noisy counterfactual. Both go into the fit.
Two experiments, one channel, and the invalid one wins
A Gaussian stand-in for the framework's in-graph calibration. The MMM's own time-series likelihood puts the channel at ROAS 1.90 ± 0.45. A geo holdout (the admissible input) measures 1.60 ± 0.35. Drag the creative A/B's reported value and its standard error and watch it overrule both. Where the earlier post's figure showed a good experiment beating a weak likelihood, this one shows two experiments competing, with the inadmissible one tighter.
At the defaults the A/B carries 8.5× the holdout's weight. The channel's calibrated ROAS moves from 1.71 to 2.88, a 68% inflation, and the 95% interval narrows from ±0.54 to ±0.22. Every downstream artifact reports a tighter number that is further from the only measurement in the room that identified the right thing.
The mechanism is the framework's code, sitting in plain sight. calibration/likelihood.py attaches an experiment with a single line whose only tuning knob is the reported standard error:
# illustrative
return pm.Normal(
name,
mu=estimand_expr, # what the model thinks the experiment measured
sigma=float(measurement.se), # how hard the measurement pulls
observed=float(measurement.value),
)
The module's own docstring says the quiet part: folding an experiment in alongside the time-series likelihood "is not double counting in the pathological sense, but the experiment's se is what governs how hard it pulls the fit." That is the right design. It is also why the admissibility question has to be answered before the payload is constructed, because after that point the only thing the model knows about your experiment is a mean and a standard deviation.
Where the Gate Belongs, and Where It Isn't
Let me be precise about what this framework does and does not currently do, because the diagnostic half of this post is a simulation and the prescriptive half is a gate that is only partly built.
First, plainly: no code in this repository models divergent delivery. Nothing estimates it, and no diagnostic would detect it. Outside the figures on this page, the framework contains no simulation of the phenomenon at all. Every occurrence of the word "divergent" in the source refers to divergent MCMC transitions or to divergent evidence across measurement sources. The audience-imbalance check that Burtch and colleagues run is a platform-side computation on user features an advertiser never receives, and an MMM cannot reconstruct it from a weekly spend panel.
Second, the place where a gate would attach is well defined, and three of its four pieces already exist. The experiment lifecycle registry in platform/sessions.py carries a design_type column per experiment and enforces a legal status path (draft, planned, running, completed, calibrated) with illegal transitions raising. The agent tool record_experiment_readout takes a method argument documented as "the measurement method (e.g. 'geo holdout DiD', 'synthetic control')". The named-method registry in planning/methods/ enumerates seven designs (synthetic control, regression-adjusted geo, TBR, GBR, matched-market DiD, ghost ads, switchback), and every one of them produces an incrementality contrast against a control or an off-state. There is no creative-A/B entry, and there could not be one under the registry's implicit admission criterion.
The fourth piece is missing, and it is the one that matters. Provenance is recorded and never read. The design_type column is unvalidated free text. The method string is written into the readout, displayed in the experiment drawer, and read by no Python module anywhere in the codebase. When apply_experiment_calibration stages a readout for the next fit, it validates exactly four things: that the channel exists in the model spec, that value and standard error are present, that a test window is present, and that out-of-window tests carry a spend level. It never asks what design produced the number. And ExperimentMeasurement, the payload that actually reaches the graph, has fifteen fields (channel, window, value, standard error, estimand, spend overrides, holdout regions, error family, off-panel evaluation settings), and not one of them describes the experiment's design.
⚠️ The evidence chip is earned by having an experiment, not by the experiment being valid
reporting/evidence.py defines the three tiers this framework prints beside every channel number: EXPERIMENT_VALIDATED, MODEL_IDENTIFIED, PRIOR_DOMINATED. The top tier's gloss describes a channel calibrated against a randomized experiment folded into the fit, then calls that "the strongest causal anchor available." Read the assignment logic and the gate is set membership: if ch in exp_set: tier = EvidenceTier.EXPERIMENT_VALIDATED. A channel calibrated from a creative A/B satisfies that test, and every client-facing surface (the classic report, the Augur readout, the interactive report) renders the strongest chip in the vocabulary next to the number that was contaminated. The tier is a claim about provenance being experimental. It is silently read as a claim about the experiment being identified.
There is one near-miss worth naming, because it shows the distinction is already half-understood in the code. reporting/triangulation.py places three estimates of a channel's return side by side (the MMM, the experiment readout, and the platform-reported figure) and tags each source with an incremental flag whose docstring calls it "the axis that makes platform-vs-MMM divergence expected, not alarming." A platform figure exceeding the incremental estimate by more than 1.4× is classified platform-inflated. That machinery would catch our 3.1× readout instantly, if it arrived labeled as a platform figure. Labeled as an experiment, it is admitted as incremental, anchors the reconciliation as the causal gold standard, and the inflation check never fires. The gate exists. The A/B walks past it wearing the wrong badge.
What Restores Validity
Four things, in ascending order of how much the practitioner will resist them.
A control cell. This is the whole fix, and it is the one that gets cut, because a holdout looks like paying for impressions you do not get and learning nothing about creative. The reframing that survives contact with a budget owner is the one this post has been building toward: the control cell is what converts an operational readout into a causal one, and a test without it produces a number that cannot enter the model. Johnson's (2023) survey of display-ad field experiments makes the same point structurally. Exposure is jointly determined by advertisers, users, algorithms and competitors, so control over exposure, not sample size, is the scarce resource.
Fixed-audience configuration. If you must run a content A/B, you can starve the algorithm of the freedom to diverge. Burtch and colleagues test this directly and report which settings help: identical optimization goals across cells, identical target audience definitions, identical budgets, identical bid strategies, manual rather than automatic placement, a single static image per cell, and a frequency cap near one impression per user. Their most restrictive combination yielded tests with no meaningful imbalance. The caveat they attach to it is the one to carry into the deck: no configuration guarantees eliminating divergent delivery entirely.
Ghost-ad controls. Where the platform supports it, the cleanest user-level design runs the campaign's auction for control users too and logs the ad they would have won without serving it (Johnson, Lewis, and Nubbemeyer, 2017). The control group is then defined by the same delivery process as the treatment group, which is exactly the property that makes the algorithm cancel. This framework's planning/methods/ghost_ads.py implements the power side of that design as a standalone pre-fit calculator. It reports the minimum detectable effect on both the intent-to-treat and the treatment-on-treated scales, since only exposure_rate of randomized users are actually reached.
A third arm that isolates the algorithm. The newest proposal in this space, and the most direct attack on the decomposition, is a three-arm design: add a cell that exposes the algorithm to the treatment's metadata while holding the user-facing creative identical to control. The difference between that arm and control is the algorithmic channel. The residual is the creative. Pal and Susarla (2026) report a Meta campaign where the algorithmic component raised female impression share by 2.07 percentage points while the creative change lowered it by 0.68, and conclude that a conventional two-arm test understates the algorithmic channel by roughly a factor of two. This is a preprint and I have not seen an independent replication. I include it because it is the only design I am aware of that targets the decomposition itself rather than suppressing it.
What This Does Not Establish
This post does not establish that platform experiments are unreliable. The evidence points the other way for the design that matters most: the Meta audit found lift-test balance indistinguishable from what randomization predicts, at the nominal 5% significance rate, across 3,204 tests. Randomized platform holdouts remain the best causal anchor an advertiser can buy, and Gordon and colleagues' (2019) large-scale comparison at Facebook is the reason this series treats them as the benchmark against which observational methods are judged. Nothing here licenses discounting a lift test.
It does not establish that A/B tests are worthless. They measure a real estimand (relative performance as the platform will actually deploy the configurations), and both papers cited here say so explicitly. A creative test that informs which asset to run next quarter on the same platform is doing legitimate work. The inadmissibility claim is scoped to one use: as a measured incrementality readout entering a model that will be used to reallocate budget across channels.
The figures are demonstrations, not estimates. The bias in figure one is generated by a one-dimensional stylized delivery model with a parameter, \( \lambda \), that has no counterpart you can read off a platform report. I do not know the magnitude of divergent-delivery bias in your account, and neither the simulation nor the audit tells you: the audit measures audience imbalance, which is an upstream symptom, and converting an observed imbalance into a bias in a conversion contrast requires exactly the responsiveness structure the simulation assumes and nobody observes. The specific claim "an SMD of 0.13 flips an 8% content difference" is arithmetic inside that model, not a measured industry number.
Finally, the framework audit is a statement about the current code, not a claim that the gate is hard to build. Three of the four pieces exist. What is missing is a validated vocabulary of design types, a read of that field in apply_experiment_calibration, and an evidence tier that distinguishes "an experiment was folded in" from "the experiment identified this estimand." That is a small, mechanical change, and until it lands, the enforcement lives in the analyst's head.
The Registry Rule
The operational version of this post fits on one line of a pre-registration form. Before an experiment's readout is allowed to become a calibration likelihood, it must name the design that produced it, and the design must have had a cell that did not receive the treatment. Geo holdout: admissible. Ghost-ad or PSA control: admissible. Switchback on-off schedule: admissible, with the block length caveats. Matched-market DiD: admissible with the usual observational caveats about parallel trends. Creative A/B, campaign-configuration A/B, audience-configuration A/B: not admissible, at any sample size, no matter how small the standard error. A small standard error makes the case for exclusion stronger, because that is when the number does the most damage.
The reporting change matters as much. A calibrated channel should print the design alongside the tier, so a reader can tell at a glance whether the strongest chip in the vocabulary was earned by a randomized holdout or by a test in which the platform's optimizer was part of the treatment. The distinction is invisible in a point estimate and a standard error, which is precisely why it has to travel as a field rather than as a footnote.
Takeaways
- Randomizing users is not sufficient to compare ad content. Delivery optimization re-selects, per creative, which randomized users are actually reached, so the measured contrast confounds content with algorithmic targeting (Braun and Schwartz, 2025).
- The asymmetry is the useful part. In a lift test both cells face the identical delivery algorithm, so it cancels. In a content A/B the algorithm is part of the treatment. Meta's own audit measured it: 0.16% of standardized mean differences above 0.20 in 3,204 lift tests versus 22% in 181,890 A/B tests, with balance tests significant at 5% and 25% respectively.
- Sample size cannot fix it. The bias is a population quantity with no \( n \) in it while the standard error falls like \( 1/\sqrt{n} \), so scale moves the test from uninformative to confidently wrong. In the simulated case that crossover lands at about 19,000 reached users per cell.
- Therefore a creative A/B result is inadmissible as an MMM calibration input. Precision-weighted pooling has no bias term: at a plausible 8.5× weight advantage the A/B moves a channel's calibrated ROAS from 1.71 to 2.88 and narrows the interval.
- In this framework the gate is three-quarters built.
design_typeandmethodare recorded but unvalidated and never read.apply_experiment_calibrationchecks channel, value, standard error and window only, andExperimentMeasurementhas no design field at all.EvidenceTier.EXPERIMENT_VALIDATEDis granted on set membership, so an inadmissible readout earns the strongest chip in the vocabulary. - What restores validity: a control cell, or failing that a fixed-audience configuration (identical goals, audiences, budgets, bids, manual placement, one static asset, frequency cap near one), ghost-ad controls, or a three-arm design that exposes the algorithm to the treatment metadata while holding the creative at control.
- This does not indict platform experiments. Lift tests came through the audit clean, and they remain the anchor. The claim is narrow, which is what makes it enforceable.
References
- Braun, M., & Schwartz, E. M. (2025). Where A/B Testing Goes Wrong: How Divergent Delivery Affects What Online Experiments Cannot (and Can) Tell You About How Customers Respond to Advertising. Journal of Marketing, 89(2), 71–95.
- Burtch, G., Moakler, R., Gordon, B. R., Zhang, P., & Hill, S. (2025). Characterizing and Minimizing Divergent Delivery in Meta Advertising Experiments. arXiv:2508.21251.
- Ali, M., Sapiezynski, P., Bogen, M., Korolova, A., Mislove, A., & Rieke, A. (2019). Discrimination Through Optimization: How Facebook's Ad Delivery Can Lead to Biased Outcomes. Proceedings of the ACM on Human-Computer Interaction, 3(CSCW), Article 199.
- Johnson, G. A. (2023). Inferno: A Guide to Field Experiments in Online Display Advertising. Journal of Economics & Management Strategy, 32(3), 469–490.
- Johnson, G. A., Lewis, R. A., & Nubbemeyer, E. I. (2017). Ghost Ads: Improving the Economics of Measuring Online Ad Effectiveness. Journal of Marketing Research, 54(6), 867–884.
- Gordon, B. R., Zettelmeyer, F., Bhargava, N., & Chapsky, D. (2019). A Comparison of Approaches to Advertising Measurement: Evidence from Big Field Experiments at Facebook. Marketing Science, 38(2), 193–225.
- Pal, P., & Susarla, A. (2026). Algorithm or Creative? A Three-Arm Experimental Design for Decomposing Algorithmic Bias in Platform A/B Tests. arXiv:2605.23706 (preprint).
- Cohen, J. (1988). Statistical Power Analysis for the Behavioral Sciences (2nd ed.). Lawrence Erlbaum Associates. (Source of the 0.20 "small effect" convention used as the balance threshold.)