A unified view of sensitivity to assumption violations
Summary
Every method in the list is the same three-step construction: (i) embed the nominal analysis in a larger family indexed by a perturbation , with the assumption as originally made; (ii) put a size on the perturbation — a hard set , a prior with a scale, a power , or an input distribution; (iii) aggregate the target over that family and report the result — worst case (robust interval, breakdown value), average (a wider posterior), derivative (local sensitivity) or variance share (Sobol index). What genuinely differs is whether the data can ever inform : for parallel trends, hidden confounding and exclusion restrictions they cannot, so the size of the perturbation must come from outside the data and “no free lunch” applies; for priors and data points they can; for ABM parameters no data are involved at all. Placebo, backdating and pre-trend checks are not sensitivity analyses — they are falsification tests, and the vault’s notes show why the two should not be confused.
Answer
1. The common skeleton
Synthesis: no single vault note states this, but every source note fits it. Let be the quantity of interest and what the analysis would report if the assumption were relaxed to . The nominal analysis reports . A sensitivity analysis chooses
- the perturbation — what is allowed to be wrong;
- a parameterisation of its size — a family with , or a distribution with scale ;
- an aggregator and a report:
| Aggregator | Formula | Report | Vault instances |
|---|---|---|---|
| Worst case | identified set, robust CI, breakdown value | Honest DiD, Rosenbaum , E-value, Cornfield | |
| Average | , as a full posterior | wider (quasi-)posterior, plotted against | Plausible GMM, Bayesian hidden- models, copula sensitivity, bias-term model |
| Derivative | at | local sensitivity diagnostic | power scaling, influence functions, OAT |
| Variance share | first-order / total-effect index | Sobol, Morris (ranking proxy) | |
| Enumeration | over a finite list | heat map, specification curve | multiverse, leave-one-out donors |
The Sensitivity Analysis in Observational Studies note calls step (i)–(ii) transparent parametrization: “explicitly separating identified parameters (which the data inform) from non-identified parameters (sensitivity parameters, which require prior information or a range of values).” That phrase is the best one-line summary of the whole family.
2. The methods side by side
| Method | (i) What is perturbed | (ii) How the size is parameterised | (iii) What is reported | Can data inform ? |
|---|---|---|---|---|
| Honest DiD | post-period differential trend | : post shocks largest pre shock; : slope changes by per period | interval “estimate minus worst-case bias”, uniform-coverage CI, breakdown | only through , as a calibration, never directly |
| Rosenbaum | treatment-assignment odds for units with identical | odds ratio | threshold where the randomization p-value crosses | no |
| E-value / Cornfield | strength of a hidden with both and | minimum risk-ratio association | one number needed to “explain away” | no |
| Hidden- logistic model (Dorie et al.); copula model (Franks et al.) | , log-odds ratios; or the dependence linking the two arms’ potential-outcome distributions | grid (frequentist) or prior (Bayesian) | over the grid, or a posterior | no |
| Plausible GMM | moment violation in (e.g. the IV exclusion restriction) | proper prior , scale multiplier | quasi-posterior for ; HPD interval plotted against | no — the prior’s impact is not asymptotically negligible |
| SC placebo / backdating / leave-one-out | which unit, which date, which donor | discrete enumeration | permutation p-value from RMSPE ratios; backdated gaps; spread of leave-one-out estimates | partly — these are tests |
| TBR stationarity stress test | stability of between pretest and test | simulated drift (+0.5%/week exponential growth), correlation | bias and coverage of the iROAS interval | partly, via cooldown flattening |
| Power scaling | the prior or the whole likelihood | exponent around 1 (e.g. 0.8, 1.25) | shift of the posterior; diagnosis; Pareto | yes — a strong likelihood swamps it |
| Data influence | one observation’s weight or value | in ; or | Pareto , pointwise elpd, influence gradient | yes |
| Sobol / Morris | model input parameters | input ranges or distributions, varied jointly | , , , | not applicable — no data involved |
| HM/ABC model discrepancy | ABM structural error | variance added to | size of the non-implausible set | weakly |
| Multiverse | the model specification itself | nodes of the model topology that pass checks | conclusions specifications heat map | yes, via model checking |
3. What is genuinely the same idea
(a) , , and the OVB product are one object. The Omitted Variables Bias formula is the prototype: bias equals (effect of the omitted thing on the outcome) (its association with treatment). Cornfield’s inequality and the E-value bound exactly those two factors (^def-e-value). The Plausible GMM IV example is the same algebra for the exclusion restriction: the exact-IV estimand equals , bias proportional to the violation and inversely proportional to first-stage strength. Honest DiD’s Lemma 2.1 is again “point estimate minus worst-case bias, given the observed pre-trend”. The Honest DiD note says so directly: ” plays the role of a Rosenbaum--type sensitivity parameter”, and the PGMM note says “plays the role of the sensitivity parameter … but here it carries a prior rather than being varied over a fixed set.”
(b) Hard set versus prior is a choice of aggregator, not of philosophy. PGMM’s Overview records that in the Gaussian limit experiment the Bayes credible interval under a two-point least-favourable prior coincides with Armstrong–Kolesár’s minimax robust interval, and that a union gives uniform coverage over a known set — the same union construction as Honest DiD’s Lemma 2.2. The Honest DiD note points the other way: a random-walk or smooth-trend prior on given in a Bayesian Difference in Differences “is the Bayesian analogue of / “. The exception is Rosenbaum’s , which the vault flags as having “no natural Bayesian analogue” because it lives inside Fisherian randomization inference.
(c) No free lunch is universal. PGMM proves it: the quasi-posterior variance is never smaller than efficient GMM’s, and efficient GMM is “the limiting, over-confident special case ” (^ex-no-free-lunch). Honest DiD’s robust sets were 40–80% longer than the estimated identified set in the VAT example, and — unlike a pre-test — they get wider when leads are noisy. Synthesis: by the Woodbury identity the plausibility-adjusted weight is : sampling variance plus misspecification variance. That is structurally the History Matching implausibility in Uncertainty Quantification for ABM Calibration, where ignoring model discrepancy “causes the HM procedure to retain too few parameter sets”. is the ABM’s . The explicit bias term , in Tail Behavior and Prior-Likelihood Conflict is the same device in a plain Bayesian model — and it shows a consequence PGMM’s Gaussian case hides: with heavy tails becomes non-monotonic in .
(d) Power scaling and data deletion are one framework. The workflow notes state that with is leave-one-out, its gradient in is an influence function, and power-scaling the prior is “the same device … applied globally” (^def-power-weighted-likelihood). PGMM’s sensitivity plot (HPD interval against prior-scale multiplier ; midpoint stable near 2.1, excluding zero throughout, ^ex-figure1) is a refit-based prior-scale analysis of exactly this kind, applied to a prior over misspecification rather than over a parameter.
(e) Static sensitivity is a Sobol index in disguise. Synthesis: static sensitivity analysis reads the scatterplot of a quantity of interest against a parameter across posterior draws — flat means insensitive to that parameter’s prior. The first-order index is the variance of the conditional-mean curve in that same scatterplot. Caveat: Sobol’s decomposition assumes independent inputs, which posterior draws are not, so this is a heuristic, not a theorem.
4. What only looks similar
- Falsification tests are not sensitivity analyses. Backdating, in-time placebos and “no significant leads” all check an observable implication in the pre-period. Pre-Trend Testing and Its Pitfalls shows the cost: a linear trend detected only 50–80% of the time can produce bias as large as the estimate, and conditioning on passing adds a further bias term . Sensitivity analysis uses the same pre-period information quantitatively instead of as a gate. Synthesis: the SC RMSPE ratio already scales post-period gaps by pre-period fit — it is a relative-magnitudes statistic — but the vault has no SC analogue of a reported breakdown . Leave-one-out over donors is closer to a true sensitivity analysis (it perturbs the estimator’s inputs), and Synthetic Control Requirements gives one directional bound: with un-excludable positive spillovers the SC estimate is a lower bound on the effect magnitude.
- Identified versus unidentified perturbations. Power scaling asks whether the data overwhelm a modelling choice; if the likelihood is strong the answer is reassuring. For , , no amount of data helps. Reporting “
priorsensefound no prior sensitivity” says nothing about confounding. - Sobol indices attribute, they do not bound. tells you which input drives output variance over an assumed input distribution; it does not say how far an input can move before a conclusion flips. It answers “which assumption should I worry about”, not “how much”.
- Local versus global cuts across all of this. OAT misses interactions ( at baseline : both inputs look irrelevant, yet ; ^oat-pitfalls). Synthesis: most causal sensitivity analyses are OAT over assumptions. The duty-to-bargain example had to move the reference period because honest sets require — assumptions interact. The multiverse is the factorial remedy, with the workflow book’s warning that “many model specifications will be nonsensical”.
Practical Implications
A checklist for any deliverable (geo test, MMM, observational lift study, ABM):
- Name the load-bearing assumption and ask whether data can ever inform it. If not (parallel trends, no hidden confounding, exclusion, TBR’s [[TBR Design Sensitivity and the Stationarity Assumption#^def-tbr-stationarity|stability of ]]), you need rows 1–6 of the table, not a diagnostic.
- Parameterise the violation in units a stakeholder can judge. Honest DiD’s marketing reading is the template: “the lift is positive unless week-to-week divergence between test and control regions during the campaign was more than times the largest divergence seen in the pre-period.” Use for weekly data with promotional shocks, for slow brand or distribution drift. For PGMM-style priors, calibrate as the AJR example does (“elasticity no larger than 10%” sd 0.05).
- Report the curve and the breakdown value, not a pass/fail. Robust interval against , HPD against , iROAS bias against drift rate. For TBR the vault only has the simulated stress test (bias and under-coverage at low ; TBR-OR unstable below ) — Synthesis: run the pseudo-geo-experiment procedure with injected drift and report the drift at which the iROAS interval first covers zero.
- Replace pre-period gates with power statements. Report what trend the pretest could have detected (Pre-Trend Testing and Its Pitfalls), and keep backdating/placebo plots as descriptive evidence.
- For the Bayesian MMM, run power scaling on the derived quantity (ROAS, optimal split), not marginal parameters — Hill parameters are individually unidentified, exactly the case where “analyzing the sensitivity of marginal posteriors is not so useful”. Sensitivity to both prior and likelihood signals conflict; do not tune priors until the warnings vanish.
- For ABMs, Morris screen then Sobol-quantify, fix inputs with , and carry explicitly; a large gap means OAT “what-if” scenarios shown to stakeholders will mislead.
- Expect to pay. If the sensitivity-aware interval is no wider than the nominal one, the perturbation set was probably set to .
Source Notes
Related Concepts
- The Selection Problem — , hidden and are all forms of selection bias being bounded rather than assumed away
- General Structure of Bayesian CI — non-identified parameters still get posteriors, so priors on them must be explicit
- Garden of Forking Paths — a multiverse chosen after seeing results is itself a forking path
- History Matching for ABMs — where enters the implausibility score
- Bayesian Structural Time-Series Model — absorbs the drift that breaks TBR, shifting the assumption rather than removing it
- Q - Uncovering Causal Estimates from Non-Experimental Data — the identification strategies whose assumptions are perturbed here
Gaps
- Partial identification proper is missing (Dream gap #30): no note on Manski’s no-assumption bounds, Balke–Pearl IV bounds or Lee trimming bounds. The vault covers the sensitivity end ( small) but not the worst-case end, so the claim that both are one continuum rests on Honest DiD alone.
- No sensitivity analysis for synthetic control or TBR that reports a breakdown value; only tests and a simulated stress test. Conformal/“honest” SC inference is not ingested.
- E-value and Rosenbaum-bound mechanics are described, not derived; regression-based tools (Cinelli–Hazlett robustness values, Oster’s ) are absent, so the OVB formula is never turned into a reportable number.
- PGMM §4 theorems (Bernstein–von Mises, coverage) are known only from the introduction, as that note’s own Gaps section records.
- Sobol with dependent inputs (Shapley effects) is not covered, which is what item 3(e) would need to be rigorous.
Follow-Up Questions
- Can a relative-magnitudes restriction be built for synthetic control, using placebo-period gaps to calibrate for the post-period gap?
- What prior on in a Bayesian DiD reproduces coverage, and how does its posterior compare with the FLCI?
- How should an MMM report a breakdown value for “unobserved demand driver correlated with spend”, combining the OVB formula with experiment-calibrated priors?