Honest DiD - Sensitivity to Parallel Trends Violations
Summary
Rambachan & Roth (2023, REStud) replace the assumption “parallel trends holds exactly” with “the post-treatment violation cannot be too different from the pre-treatment violations ,” formalised as for a researcher-chosen set. Two leading choices: relative magnitudes (post-period shocks at most times the largest pre-period shock) and smoothness (the differential trend’s slope changes by at most per period; is a linear-trend extrapolation). Under such restrictions the treatment effect is partially identified; the paper supplies confidence sets with uniform coverage that account for both sampling error in the estimated pre-trend and identification uncertainty. The natural deliverable is a sensitivity plot and a breakdown value — the smallest (or ) at which the conclusion no longer holds. Implemented in the
HonestDiDR/Stata packages.
Overview
Pre-Trend Testing and Its Pitfalls ends with a problem: the pre-test is underpowered, distorts inference when passed, and gives no guidance when failed. Honest DiD keeps the intuition behind pre-testing — pre-trends are informative about counterfactual post-trends — but uses it quantitatively rather than as a gate. Two kinds of uncertainty are separated (§1):
- Statistical uncertainty — “we can only noisily estimate the true pre-trend.”
- Identification uncertainty — “even if the true pre-trend were known, we may not know exactly how to extrapolate it.”
A useful consequence noted by Roth et al. (2023, §4.5): unlike a pre-test, honest confidence sets get wider when the leads are imprecisely estimated. A noisy, “insignificant” pre-period is no longer rewarded.
The approach builds on Manski & Pepper (2018), who bounded DiD variation in their right-to-carry application by calibrating to the largest pre-period deviation; Rambachan & Roth generalise the restriction classes and, crucially, add inference.
Main Content
Set-up and identified set ^def-honest-identified-set
Event-study coefficients with and , (decomposition). Target: a scalar (one period’s effect, or the post-period average with ). Assume . The identified set is
Exact parallel trends is the special case .
Lemma 2.1 — the identified set is an interval ^thm-honest-lemma21
If is closed and convex, with
i.e. “point estimate minus worst-case bias, given the observed pre-trend.” If , then , and a union of valid confidence sets is valid (Lemma 2.2).
Restriction classes (§2.4)
With at the reference period:
Relative magnitudes
Motivation: confounding shocks post-treatment are of similar size to those seen pre-treatment. is the natural benchmark when pre- and post-windows have similar length.
Smoothness (second differences)
Motivation: smoothly evolving secular trends. forces to be exactly linear — the assumption behind adding group-specific linear trends — so nests and relaxes that common practice.
Hybrids and shape restrictions: bounds post-period non-linearity by times the largest pre-period non-linearity; sign (: for , e.g. a known confounding policy with positive effect) and monotonicity () restrictions can be intersected with the above.
Polyhedral class (Definition 1): . is a polyhedron; and are finite unions of polyhedra (one per location and sign of the pre-period maximum). Bespoke restrictions — e.g. Ashenfelter’s dip — fit the same template.
Three-period illustration (): under , , so the identified set for is . Under , , so the set is : extrapolate the pre-trend linearly, then allow slack .
Inferential goal — uniform ("honest") coverage ^def-honest-coverage
Under the normal approximation , find with
(eq. 10), which translates to uniform asymptotic coverage over a large class of DGPs with estimated (eq. 11). Only and its full covariance matrix are required — any estimator that delivers them (TWFE, Callaway–Sant’Anna, Sun–Abraham, IV event studies) can be plugged in.
Method 1 — moment inequalities (conditional / hybrid tests)
For polyhedral , since , the null is equivalent to a system of moment inequalities with linear nuisance parameters (eq. 13):
The nuisance parameters are the other post-period effects. The test statistic profiles them out with a linear program, s.t. (eq. 14), whose dual is s.t. . Following Andrews, Roth & Pakes (ARP), conditional on the optimal dual vertex , is truncated normal, giving a critical value that adapts to which inequalities bind. The hybrid test first runs a size- least-favourable test, then a conditional test of adjusted size, fixing the conditional test’s weak power when two moments are nearly tied. Confidence sets come from test inversion over . Rambachan & Roth prove uniform size control, consistency, and — under a linear-independence constraint qualification (LICQ) — optimal local asymptotic power. The approach is tractable even when (true of 5 of the 12 papers in Roth’s survey).
Method 2 — fixed-length confidence intervals (FLCIs)
Optimal FLCI and its guarantees (§4) ^thm-honest-flci
Consider . For given the smallest valid half-length is
where is the worst-case bias of the affine estimator over (eq. 17) and is the quantile of . Minimise over (a convex problem when is convex). For and the optimal estimator is
— the first lag minus a weighted average of pre-period slopes, trading bias (favours recent slopes) against variance (favours averaging many).
Proposition 4.1: if is convex and centrosymmetric (true for ; false for , ) and is, e.g., zero or linear, the optimal FLCI’s length is near-minimal in finite samples: at any valid confidence set is at most 28% shorter in expectation.
Proposition 4.2: an FLCI is consistent iff the identified set has its maximal possible length at the true . For with every affine estimator has infinite worst-case bias, so “the only valid FLCI is the entire real line”; FLCIs also ignore added sign/monotonicity restrictions.
Recommendation (§5.2.5, §6.1.2): FLCIs for ; the ARP hybrid for everything else, including . HonestDiD applies these defaults.
Choosing and reporting
- Choose the class from domain knowledge about the confounder: discrete differential shocks ; smooth secular trends ; known sign of a concurrent confounder add .
- Report the robust confidence set as a function of (or ), and the breakdown value. If is the population breakdown point for a null effect and its sample analogue, is a valid confidence interval for (fn. 33).
- Identified sets under grow with the horizon — per-period deviations accumulate — so average or late-period effects are less robust than first-period effects.
- Interpret with context: “robust to ” is strong if the treatment date was otherwise quiet, weak if it coincided with a shock larger than anything in the pre-period (Roth et al. §4.6).
- Later pre-periods can be given more weight, e.g. (fn. 7).
Examples
Benzarti & Carloni — French restaurant VAT cut (§6.2) ^ex-honest-vat
Event study of log restaurant profits vs other service firms, VAT cut July 2009. The pre-test fails: is rejected (). Industry-specific shocks, not smooth trends, are the concern, so use . For the 2009 effect, gives a robust confidence set of — wider than OLS but excluding zero. The breakdown value is : the conclusion survives unless post-period shocks were more than twice the largest pre-period shock (2009 was a recession, so this is a live question). For the four-year average effect, the set already includes zero and is about twice as wide. Confidence sets are 40–80% longer than the estimated identified set — both sources of uncertainty matter.
Lovenheim & Willén — duty-to-bargain laws (§6.3) ^ex-honest-dtb
Long-run employment effects, concern = smooth secular trends, so with FLCIs for . Men: robust sets resemble OLS near ; breakdown at . Women: OLS is significantly negative, but a visible downward pre-trend means that for the robust set contains only positive values — the point estimate lies above the linear extrapolation of the pre-trend. Calibration: equals the employment effect of s.d. of teacher value-added per period. Also note the reference period was moved from to because cohorts at may be partially treated — honest sets require .
Workflow sketch in R:
library(did); library(HonestDiD)
es <- aggte(att_gt(yname="y", tname="t", idname="id", gname="g", data=df, base_period="universal"),
type = "dynamic") # heterogeneity-robust event study
# betahat: event-study coefficients (reference period removed); sigma: their FULL covariance
# (for `did` output, build sigma from the influence functions; see the pedrohcgs/CS_RR example repo)
rm <- createSensitivityResults_relativeMagnitudes(betahat, sigma,
numPrePeriods = K, numPostPeriods = M_post, Mbarvec = seq(0.5, 2, by = 0.5))
sd <- createSensitivityResults(betahat, sigma,
numPrePeriods = K, numPostPeriods = M_post, Mvec = seq(0, 0.05, by = 0.01))
orig <- constructOriginalCS(betahat, sigma, numPrePeriods = K, numPostPeriods = M_post)
createSensitivityPlot_relativeMagnitudes(rm, orig) # read off the breakdown MbarMarketing reading. For an observational lift study (regional price change, staggered feature rollout), the stakeholder-facing statement becomes: “the lift is positive unless week-to-week divergence between test and control regions during the campaign was more than times the largest divergence seen in the pre-period.” With weekly data and promotional shocks, is usually the more defensible class; with slow distribution or brand-health drift, .
Connections
- Pre-Trend Testing and Its Pitfalls — the problem this solves.
- Event Study Designs and Dynamic Treatment Effects — supplies and ; any asymptotically normal event-study estimator qualifies.
- Aggregating Group-Time Effects, Group-Time Average Treatment Effects — honest sets can be wrapped around Callaway–Sant’Anna event-study aggregates.
- Sensitivity Analysis in Observational Studies — same philosophy (bound the unverifiable, report a breakdown point) applied to unobserved confounding under ignorability; plays the role of a Rosenbaum--type sensitivity parameter.
- Plausible GMM - Overview — a quasi-Bayesian relative: instead of a hard set for the violation, place a prior on the degree of moment misspecification.
- Bayesian Difference in Differences — a prior over given (e.g. a random walk or smooth-trend prior) is the Bayesian analogue of / .
- Synthetic Difference-in-Differences - Overview — the alternative strategy: construct parallel trends by reweighting. The two are complementary, not competing.
- The Selection Problem — is the DiD form of selection bias; Honest DiD bounds it, in the partial-identification tradition of Manski, rather than assuming it away.