A unified view of sensitivity to assumption violations

Summary

Every method in the list is the same three-step construction: (i) embed the nominal analysis in a larger family indexed by a perturbation , with the assumption as originally made; (ii) put a size on the perturbation — a hard set , a prior with a scale, a power , or an input distribution; (iii) aggregate the target over that family and report the result — worst case (robust interval, breakdown value), average (a wider posterior), derivative (local sensitivity) or variance share (Sobol index). What genuinely differs is whether the data can ever inform : for parallel trends, hidden confounding and exclusion restrictions they cannot, so the size of the perturbation must come from outside the data and “no free lunch” applies; for priors and data points they can; for ABM parameters no data are involved at all. Placebo, backdating and pre-trend checks are not sensitivity analyses — they are falsification tests, and the vault’s notes show why the two should not be confused.

Answer

1. The common skeleton

Synthesis: no single vault note states this, but every source note fits it. Let be the quantity of interest and what the analysis would report if the assumption were relaxed to . The nominal analysis reports . A sensitivity analysis chooses

  1. the perturbation — what is allowed to be wrong;
  2. a parameterisation of its size — a family with , or a distribution with scale ;
  3. an aggregator and a report:
AggregatorFormulaReportVault instances
Worst caseidentified set, robust CI, breakdown value Honest DiD, Rosenbaum , E-value, Cornfield
Average, as a full posteriorwider (quasi-)posterior, plotted against Plausible GMM, Bayesian hidden- models, copula sensitivity, bias-term model
Derivative at local sensitivity diagnosticpower scaling, influence functions, OAT
Variance sharefirst-order / total-effect indexSobol, Morris (ranking proxy)
Enumeration over a finite listheat map, specification curvemultiverse, leave-one-out donors

The Sensitivity Analysis in Observational Studies note calls step (i)–(ii) transparent parametrization: “explicitly separating identified parameters (which the data inform) from non-identified parameters (sensitivity parameters, which require prior information or a range of values).” That phrase is the best one-line summary of the whole family.

2. The methods side by side

Method(i) What is perturbed(ii) How the size is parameterised(iii) What is reportedCan data inform ?
Honest DiDpost-period differential trend : post shocks largest pre shock; : slope changes by per periodinterval “estimate minus worst-case bias”, uniform-coverage CI, breakdown only through , as a calibration, never directly
Rosenbaum treatment-assignment odds for units with identical odds ratio threshold where the randomization p-value crosses no
E-value / Cornfieldstrength of a hidden with both and minimum risk-ratio associationone number needed to “explain away”no
Hidden- logistic model (Dorie et al.); copula model (Franks et al.), log-odds ratios; or the dependence linking the two arms’ potential-outcome distributionsgrid (frequentist) or prior (Bayesian) over the grid, or a posteriorno
Plausible GMMmoment violation in (e.g. the IV exclusion restriction)proper prior , scale multiplier quasi-posterior for ; HPD interval plotted against no — the prior’s impact is not asymptotically negligible
SC placebo / backdating / leave-one-outwhich unit, which date, which donordiscrete enumerationpermutation p-value from RMSPE ratios; backdated gaps; spread of leave-one-out estimatespartly — these are tests
TBR stationarity stress teststability of between pretest and testsimulated drift (+0.5%/week exponential growth), correlation bias and coverage of the iROAS intervalpartly, via cooldown flattening
Power scalingthe prior or the whole likelihoodexponent around 1 (e.g. 0.8, 1.25)shift of the posterior; diagnosis; Pareto yes — a strong likelihood swamps it
Data influenceone observation’s weight or value in ; or Pareto , pointwise elpd, influence gradientyes
Sobol / Morrismodel input parameters input ranges or distributions, varied jointly, , , not applicable — no data involved
HM/ABC model discrepancyABM structural errorvariance added to size of the non-implausible setweakly
Multiversethe model specification itselfnodes of the model topology that pass checksconclusions specifications heat mapyes, via model checking

3. What is genuinely the same idea

(a) , , and the OVB product are one object. The Omitted Variables Bias formula is the prototype: bias equals (effect of the omitted thing on the outcome) (its association with treatment). Cornfield’s inequality and the E-value bound exactly those two factors (^def-e-value). The Plausible GMM IV example is the same algebra for the exclusion restriction: the exact-IV estimand equals , bias proportional to the violation and inversely proportional to first-stage strength. Honest DiD’s Lemma 2.1 is again “point estimate minus worst-case bias, given the observed pre-trend”. The Honest DiD note says so directly: ” plays the role of a Rosenbaum--type sensitivity parameter”, and the PGMM note says “plays the role of the sensitivity parameter … but here it carries a prior rather than being varied over a fixed set.”

(b) Hard set versus prior is a choice of aggregator, not of philosophy. PGMM’s Overview records that in the Gaussian limit experiment the Bayes credible interval under a two-point least-favourable prior coincides with Armstrong–Kolesár’s minimax robust interval, and that a union gives uniform coverage over a known set — the same union construction as Honest DiD’s Lemma 2.2. The Honest DiD note points the other way: a random-walk or smooth-trend prior on given in a Bayesian Difference in Differences “is the Bayesian analogue of / “. The exception is Rosenbaum’s , which the vault flags as having “no natural Bayesian analogue” because it lives inside Fisherian randomization inference.

(c) No free lunch is universal. PGMM proves it: the quasi-posterior variance is never smaller than efficient GMM’s, and efficient GMM is “the limiting, over-confident special case ” (^ex-no-free-lunch). Honest DiD’s robust sets were 40–80% longer than the estimated identified set in the VAT example, and — unlike a pre-test — they get wider when leads are noisy. Synthesis: by the Woodbury identity the plausibility-adjusted weight is : sampling variance plus misspecification variance. That is structurally the History Matching implausibility in Uncertainty Quantification for ABM Calibration, where ignoring model discrepancy “causes the HM procedure to retain too few parameter sets”. is the ABM’s . The explicit bias term , in Tail Behavior and Prior-Likelihood Conflict is the same device in a plain Bayesian model — and it shows a consequence PGMM’s Gaussian case hides: with heavy tails becomes non-monotonic in .

(d) Power scaling and data deletion are one framework. The workflow notes state that with is leave-one-out, its gradient in is an influence function, and power-scaling the prior is “the same device … applied globally” (^def-power-weighted-likelihood). PGMM’s sensitivity plot (HPD interval against prior-scale multiplier ; midpoint stable near 2.1, excluding zero throughout, ^ex-figure1) is a refit-based prior-scale analysis of exactly this kind, applied to a prior over misspecification rather than over a parameter.

(e) Static sensitivity is a Sobol index in disguise. Synthesis: static sensitivity analysis reads the scatterplot of a quantity of interest against a parameter across posterior draws — flat means insensitive to that parameter’s prior. The first-order index is the variance of the conditional-mean curve in that same scatterplot. Caveat: Sobol’s decomposition assumes independent inputs, which posterior draws are not, so this is a heuristic, not a theorem.

4. What only looks similar

  1. Falsification tests are not sensitivity analyses. Backdating, in-time placebos and “no significant leads” all check an observable implication in the pre-period. Pre-Trend Testing and Its Pitfalls shows the cost: a linear trend detected only 50–80% of the time can produce bias as large as the estimate, and conditioning on passing adds a further bias term . Sensitivity analysis uses the same pre-period information quantitatively instead of as a gate. Synthesis: the SC RMSPE ratio already scales post-period gaps by pre-period fit — it is a relative-magnitudes statistic — but the vault has no SC analogue of a reported breakdown . Leave-one-out over donors is closer to a true sensitivity analysis (it perturbs the estimator’s inputs), and Synthetic Control Requirements gives one directional bound: with un-excludable positive spillovers the SC estimate is a lower bound on the effect magnitude.
  2. Identified versus unidentified perturbations. Power scaling asks whether the data overwhelm a modelling choice; if the likelihood is strong the answer is reassuring. For , , no amount of data helps. Reporting “priorsense found no prior sensitivity” says nothing about confounding.
  3. Sobol indices attribute, they do not bound. tells you which input drives output variance over an assumed input distribution; it does not say how far an input can move before a conclusion flips. It answers “which assumption should I worry about”, not “how much”.
  4. Local versus global cuts across all of this. OAT misses interactions ( at baseline : both inputs look irrelevant, yet ; ^oat-pitfalls). Synthesis: most causal sensitivity analyses are OAT over assumptions. The duty-to-bargain example had to move the reference period because honest sets require — assumptions interact. The multiverse is the factorial remedy, with the workflow book’s warning that “many model specifications will be nonsensical”.

Practical Implications

A checklist for any deliverable (geo test, MMM, observational lift study, ABM):

  1. Name the load-bearing assumption and ask whether data can ever inform it. If not (parallel trends, no hidden confounding, exclusion, TBR’s [[TBR Design Sensitivity and the Stationarity Assumption#^def-tbr-stationarity|stability of ]]), you need rows 1–6 of the table, not a diagnostic.
  2. Parameterise the violation in units a stakeholder can judge. Honest DiD’s marketing reading is the template: “the lift is positive unless week-to-week divergence between test and control regions during the campaign was more than times the largest divergence seen in the pre-period.” Use for weekly data with promotional shocks, for slow brand or distribution drift. For PGMM-style priors, calibrate as the AJR example does (“elasticity no larger than 10%” sd 0.05).
  3. Report the curve and the breakdown value, not a pass/fail. Robust interval against , HPD against , iROAS bias against drift rate. For TBR the vault only has the simulated stress test (bias and under-coverage at low ; TBR-OR unstable below ) — Synthesis: run the pseudo-geo-experiment procedure with injected drift and report the drift at which the iROAS interval first covers zero.
  4. Replace pre-period gates with power statements. Report what trend the pretest could have detected (Pre-Trend Testing and Its Pitfalls), and keep backdating/placebo plots as descriptive evidence.
  5. For the Bayesian MMM, run power scaling on the derived quantity (ROAS, optimal split), not marginal parameters — Hill parameters are individually unidentified, exactly the case where “analyzing the sensitivity of marginal posteriors is not so useful”. Sensitivity to both prior and likelihood signals conflict; do not tune priors until the warnings vanish.
  6. For ABMs, Morris screen then Sobol-quantify, fix inputs with , and carry explicitly; a large gap means OAT “what-if” scenarios shown to stakeholders will mislead.
  7. Expect to pay. If the sensitivity-aware interval is no wider than the nominal one, the perturbation set was probably set to .

Source Notes

NoteRelevance
Honest DiD - Sensitivity to Parallel Trends Violations, , Lemma 2.1, breakdown value, marketing reading
Pre-Trend Testing and Its PitfallsWhy tests of assumptions are not sensitivity analyses
Sensitivity Analysis in Observational StudiesCornfield, E-value, , copula; “transparent parametrization”
Plausible GMM - Overview, Plausible Moment Restriction Model, Quasi-Bayes for Plausible Moment Restrictions, Gaussian Local Prior Approximation, Plausible GMM - Institutions and GDP ApplicationPrior over misspecification, no free lunch, link to minimax intervals, HPD-vs-scale plot
Synthetic Control Inference and Diagnostics, Synthetic Control RequirementsRMSPE ratio, backdating, leave-one-out, spillover lower bound
TBR Design Sensitivity and the Stationarity AssumptionSimulated violation of stationarity, TBR-OR
Influence of Likelihood and Prior, Influence of Individual Data Points, Tail Behavior and Prior-Likelihood ConflictPower scaling, static sensitivity, moving vs removing, bias-term model
Global Sensitivity Analysis - Overview, Variance-Based Sensitivity and Sobol Indices, Morris Elementary Effects Screening, Local vs Global Sensitivity AnalysisVariance-share aggregator, OAT failure
Uncertainty Quantification for ABM CalibrationModel discrepancy
Omitted Variables Bias, Instrumental VariablesThe bias algebra being bounded
Topology of Models, Comparing Models VisuallyMultiverse as enumeration over specifications

Gaps

  • Partial identification proper is missing (Dream gap #30): no note on Manski’s no-assumption bounds, Balke–Pearl IV bounds or Lee trimming bounds. The vault covers the sensitivity end ( small) but not the worst-case end, so the claim that both are one continuum rests on Honest DiD alone.
  • No sensitivity analysis for synthetic control or TBR that reports a breakdown value; only tests and a simulated stress test. Conformal/“honest” SC inference is not ingested.
  • E-value and Rosenbaum-bound mechanics are described, not derived; regression-based tools (Cinelli–Hazlett robustness values, Oster’s ) are absent, so the OVB formula is never turned into a reportable number.
  • PGMM §4 theorems (Bernstein–von Mises, coverage) are known only from the introduction, as that note’s own Gaps section records.
  • Sobol with dependent inputs (Shapley effects) is not covered, which is what item 3(e) would need to be rigorous.

Follow-Up Questions

  • Can a relative-magnitudes restriction be built for synthetic control, using placebo-period gaps to calibrate for the post-period gap?
  • What prior on in a Bayesian DiD reproduces coverage, and how does its posterior compare with the FLCI?
  • How should an MMM report a breakdown value for “unobserved demand driver correlated with spend”, combining the OVB formula with experiment-calibrated priors?