Doubly-robust estimation appears in the vault as AIPW, the DML interactive-model score, the Callaway–Sant’Anna doubly-robust ATT(g,t), SDID’s double robustness, the X-learner and weighted conformal prediction. What is the common structure, and in what sense is each one “doubly” robust?

Summary

Every genuinely doubly-robust procedure in the vault has the same skeleton: a model-based prediction plus a weighted average of that model’s residuals, arranged so the leading error is a product (error of the outcome-type nuisance) (error of the weight-type nuisance). “Doubly” then means one of three different things: model DR (consistent if either nuisance is correctly specified — AIPW, Callaway–Sant’Anna, A-learning), rate DR (-normal if the product of the two ML error rates is — DML, and in function-valued form the R-learner and locally-centered GRF), or SDID’s balancing DR (bias vanishes if either the unit weights or the time weights balance the latent factors). Weighted conformal prediction is doubly robust for coverage, by a different mechanism; the X-learner is not doubly robust at all, despite using both outcome models and a propensity score; and a dogmatically Bayesian analysis has no DR analogue, because the propensity score drops out of the likelihood.

Answer

1. The shared skeleton

The template is the AIPW estimator of ^def-dr: with . DML Estimators for ATE and the Interactive Model writes the same thing as a score,

and Neyman Orthogonality explains where it comes from: “orthogonal score = original score + influence-function adjustment” — the plug-in plus the correction for having estimated .

Synthesis (algebra not written out in any one note, but implied by the DML note’s “the second-order term is a cross-product ”): evaluate the treated-arm part of the score at arbitrary fixed and take expectations using :

That single product is the whole story. Read it three ways:

Reading of the product NameStatement in the vault
Either factor is identically zeroModel double robustness”consistent if either the propensity score model or the outcome model is correctly specified” (Frequentist Causal Estimation)
Its derivative in each factor is zero at the truthNeyman orthogonality; first-order insensitivity to both nuisances (Neyman Orthogonality)
Both factors shrink and the product is Rate double robustness (Theorem 5.1 of the DML note)

Neyman Orthogonality says it directly: double robustness and orthogonality “are two faces of the same product-form remainder.”

2. Model DR versus rate DR

These are different guarantees for different workflows.

  • Model DR is a statement about parametric misspecification: fit a logit and an OLS, one may be wrong forever, and the estimator is still consistent. It says nothing about the rate or the standard errors when one model is wrong.
  • Rate DR is a statement about regularised ML learners that are both right in the limit but slow. With both nuisances at and , a non-orthogonal score carries bias of about four standard-error units, the orthogonal one about (Neyman Orthogonality, numerical intuition). The trade-off is explicit in Remark 5.2: with sparsity indices the requirement is , “much weaker than ” — a very sparse propensity buys a dense outcome regression and vice versa. It needs cross-fitting as the second ingredient, and delivers more than consistency: uniform -normality and Hahn’s efficiency bound.
  • Known design collapses both. When is known (an RCT), the second derivative vanishes and “only consistency of is needed”; the outcome model becomes pure variance reduction (Pennsylvania bonus example, where every learner gives to ).

3. Each appearance, side by side

AppearanceOutcome-type nuisanceWeight-type nuisanceSense of “doubly”What is protected
AIPW (Frequentist Causal Estimation)Model DRConsistency of the PATE
DML interactive model (DML Estimators for ATE and the Interactive Model)Rate DR (product rate) + orthogonality-normal, efficient ATE / ATTE / LATE
Callaway–Sant’Anna (Doubly-Robust Estimands for ATT(g,t))generalized propensity Model DR (logit + OLS in practice)Consistency of each under conditional parallel trends
A-learning (A-learning and Robustness)baseline propensity Model DR, conditional on a correct contrast Consistency of the optimal-regime parameters
R-learner, GRF local centering (R-Learner and Orthogonal CATE Estimation, Generalized Random Forests - Local Moment Equations)Rate-type (quasi-oracle): both at Oracle regret for the function
SDID (SDID vs DiD vs Synthetic Control)time weights (a regression of post on pre)unit weights (a regression of treated on controls)Balancing DR over latent factorsBias under
Weighted split-CQR (Conformal Inference for Counterfactuals and ITEs)conditional quantiles likelihood-ratio weightsEither-or, for coverage
X-learner (X-Learner) as mixing weightNone—

Notes on the rows that differ from plain AIPW:

Callaway–Sant’Anna. Theorem 1 shows OR, IPW and DR estimands are identical as identification targets and differ only once nuisances are fitted. The DR estimand is literally (treated weight propensity-odds comparison weight) (long difference outcome regression) — weights times residuals. The DML note identifies its ATTE score (5.4) as the cross-sectional analogue, which is what would license ML nuisances. What is made robust is covariate adjustment inside parallel trends, not parallel trends itself.

SDID. The bias can be grouped as a unit-weighted contrast of time-regression residuals or a time-weighted contrast of unit-regression residuals, so it vanishes if either regression “fits and generalises”; “even if neither model generalizes sufficiently well on its own, it suffices for one model to predict the generalization error of the other.” The authors themselves liken this to AIPW, and the equivalence with augmented synthetic control (linear ) shows the mapping: time weights play the outcome model, unit weights play the propensity/balancing weights (Synthetic Control Extensions gives the same “SC applied to regression residuals” form). What is different: the nuisances are functions of a latent matrix , not of observed ; there is no correct-specification statement, only an untestable one (“essentially an assumption of no unexplained confounding”); and the asymptotics need .

Weighted conformal. Theorem 1 of Lei & Candès is an either-or result (A1: consistent; A2: quantiles consistent), but the mechanism is not a product-form bias. If the weights are right, Proposition 1 gives coverage “uniform over all and all quantile learners”; if the quantiles are right, is approximately the quantile of any reweighting of the scores, so the weights stop mattering. With wrong weights and arbitrary quantiles the loss is bounded by . The authors stress that “consistency of point estimates and coverage of interval estimates are different concepts.” The weights are the same propensity odds as in IPW (Conformal Prediction Under Covariate Shift, Propensity Score and the Balancing Property); the object protected is different.

X-learner — looks similar, is not. It fits both outcome models, cross-imputes and , and combines with . The propensity here is a variance-driven mixing weight (“the better-estimated CATE component dominates”), not a residual reweighting, and the vault’s X-learner note never claims double robustness — its guarantee is a minimax rate . R-Learner and Orthogonal CATE Estimation records Nie & Wager’s counterexample: shift down and up by and moves by exactly that amount — first-order sensitivity. Synthesis: this is immediate from the combination rule, since the shift contributes whatever is. The doubly-robust CATE estimators in the vault are the R-learner and locally-centered GRF; the “DR-learner” is mentioned once in Frequentist Causal Estimation but has no note.

4. Is there a Bayesian analogue?

Strictly, no. Under ignorability and prior independence the assignment model factors out, so “the Bayesian posterior for causal estimands depends only on the outcome model” (Propensity Score in Bayesian CI); Robins, Hernán & Wasserman: “Bayesian inference must ignore the propensity score” (Bayesian Inverse Probability Weighting). A pure Bayesian analysis is therefore singly robust by construction, and the failure mode has a Bayesian name — regularization-induced confounding, where shrinkage priors on nuisance coefficients “effectively remove confounding regardless of what the data say” (^warn-reg-confounding); BART is also “overconfident in poor overlap regions” (^ex-41).

What the vault offers instead:

DeviceWhat it isDR status
Liao–Zigler Bayesian IPWPosterior draws of , weighted outcome fit per draw, Rubin’s rules (SE vs naive )Not DR. Propagates propensity uncertainty; still singly robust on the propensity side, and “only quasi-Bayesian”
as a covariate, ; BCFOutcome regression on propensity strataCalled “Bayesian double robustness” (Wang 2012; Saarela 2016): redundant if the outcome model is right, balancing if it is wrong. A robustness heuristic — no product-rate theorem; two-stage, and joint fitting has the feedback problem
Dependent priorsLink a priori; Zigler–Dominici recovers Hájek IPW as a posterior meanBayesian justification of IPW, not DR; “no general recipe”
Posterior-predictive plug-in (Ding & Liu)Draw both models, push draws through Inherits frequentist DR; “not dogmatically Bayesian”
Bayesian bootstrapDR as an M-estimation problem under a Dirichlet-process prior (General Structure of Bayesian CI)DR with Bayesian uncertainty, but gives up informative priors

So the honest summary: the Bayesian analogue of double robustness is design-stage use of the propensity score plus a flexible outcome prior, or a hybrid that feeds posterior draws into a frequentist DR functional. None of these makes the posterior itself doubly robust.

5. What “doubly robust” never buys

  1. Identification. Unconfoundedness, conditional parallel trends, or SDID’s no-unexplained-confounding are assumed, not protected: “ML on cannot fix a missing confounder.”
  2. Overlap. enters the score; “a highly predictive is not ‘good’: it signals limited overlap” (Common Support and Overlap). Conformal is the one method that fails loudly — the interval becomes where .
  3. Both wrong. No guarantee; in the A-learning simulations “neither dominates uniformly.”
  4. Free efficiency. A-learning pays an efficiency price when everything is right; SDID’s unequal weights “may worsen the precision” on pure-noise panels.

Practical Implications

  • User-level ad experiments (known randomisation). Use cross-fit AIPW with the design propensity. Any consistent outcome learner is enough, the CI is , and weighted-CQR ITE intervals are finite-sample exact. Here DR is a precision device, not a bias device.
  • Observational exposure data (ad targeting). Rate DR is the relevant notion: cross-fit, clip (the DML sketch uses ), and read a near-perfect exposure classifier as a warning about overlap rather than a success. Report how much the estimate moves across learners, as in the 401(k) table.
  • Staggered geo roll-outs with covariates. Callaway–Sant’Anna DR per , then aggregate. Remember its robustness is over covariate adjustment only.
  • Hand-picked geo tests. SDID’s double robustness is over latent market factors; it is why SDID is roughly unbiased where DiD is “visibly off-centre”, and under random assignment it is simply more precise (RMSE vs ).
  • Bayesian MMM. Synthesis: an MMM is an outcome-model-only analysis, so it sits in the singly-robust column and is exposed to regularization-induced confounding whenever shrinkage priors are put on control variables that also drive spend. The closest available safeguards are (i) residualising spend on its drivers as a sensitivity fit — the PLR/GRF local-centering moment accepts continuous treatments — and (ii) anchoring with experiments, where the assignment mechanism is known.
  • Checklist. (a) Which two nuisances? (b) Is the remainder a product of their errors, or is one of them merely a mixing weight (X-learner)? (c) Parametric fits → claim model DR; ML fits → cross-fit and claim rate DR. (d) Is the design known? Then the outcome model only buys precision. (e) What is protected — a point estimate, a function, or coverage? (f) Check overlap before anything else.

Source Notes

NoteRelevance
Frequentist Causal EstimationAIPW definition; either-or theorem
DML Estimators for ATE and the Interactive Model, Neyman OrthogonalityAIPW as orthogonal score; product-rate condition; known-propensity case
Doubly-Robust Estimands for ATT(g,t)OR / IPW / DR estimands for staggered DiD
SDID vs DiD vs Synthetic Control, Synthetic Control ExtensionsTwo groupings of the bias; ASCM equivalence; untestable-assumption caveat
A-learning and RobustnessDR conditional on a correct contrast; efficiency price
X-Learner, R-Learner and Orthogonal CATE EstimationPropensity as mixing weight; quasi-oracle property; counterexample
Generalized Random Forests - Local Moment EquationsLocal centering; continuous treatments
Conformal Inference for Counterfactuals and ITEs, Conformal Prediction Under Covariate ShiftDouble robustness of coverage; bound; propensity-odds weights
Bayesian Inverse Probability Weighting, Propensity Score in Bayesian CIWhy weights are not in the likelihood; Liao–Zigler; three strategies; feedback problem
Bayesian Outcome Models, General Structure of Bayesian CIRegularization-induced confounding; BCF; Bayesian bootstrap
Propensity Score and the Balancing Property, Common Support and OverlapThe weight-side nuisance and positivity
Chernozhukov 2018 - Double Debiased Machine Learning§5.1, Theorem 5.1, Remark 5.2
Arkhangelsky 2021 - Synthetic Difference in Differences§4.2, p. 23
Lei Candes 2020 - Conformal Inference of Counterfactuals and ITEsProposition 1, Theorem 1
Li et al. - 2022 - Bayesian causal inference a critical review§5, pp. 10–13

Gaps

  • No DR-learner note. Frequentist Causal Estimation points to a “DR-learner” in Metalearners for CATE, but that note covers only S-, T- and X-learners. The pseudo-outcome regression that is the doubly-robust CATE metalearner is missing.
  • No TMLE or general semiparametric-efficiency note; efficient influence functions appear only through Neyman Orthogonality.
  • Continuous-treatment DR (generalized propensity score, dose-response) is absent — the case that matters for media spend. Only the PLR/GRF partialling-out moment covers continuous .
  • Bayesian DR is second-hand: Wang et al., Saarela et al., Ding & Liu and the Bayesian bootstrap are one-paragraph summaries from the Li et al. review; no worked example.
  • Both-models-wrong behaviour (the Kang & Schafer debate) is only name-checked in the conformal note.
  • Conformal DR: the vault records the either-or theorem and the bound but no product-type bound, so whether coverage error is second order is not answered here.
  • SDID: the formal conditions under which “one regression predicts the generalization error of the other” are stated only informally.

Follow-Up Questions

  • What does the DR-learner pseudo-outcome look like, and how does it compare with the R-learner under weak overlap?
  • Can a geo-level Bayesian MMM be given a design-stage “spend propensity” in the spirit of BCF’s covariate, and does it reduce regularization-induced confounding?
  • How do Callaway–Sant’Anna DR estimates change when the logit/OLS nuisances are replaced by cross-fit ML?
  • Is there a doubly-robust correction for a simulator: ABM prediction plus weighted residuals from experimental cells?