Pre-Trend Testing and Its Pitfalls
Summary
The standard defence of parallel trends — “no pre-treatment event-study coefficient is significant” — fails in two distinct ways (Roth 2022). Low power: in simulations calibrated to 12 published AEA-journal event studies, a linear differential trend that the pre-test catches only 50–80% of the time often produces bias as large as the estimated treatment effect, and nominal 95% intervals that exclude the truth up to 98% of the time. Pre-test bias: the datasets that pass the pre-test are a selected sample. Conditional on passing, , and under homoskedasticity with a monotone trend the extra term has the same sign as the trend bias — passing the test makes things worse. The prescription is not to stop plotting event studies but to report power against relevant violations and to replace the binary test with sensitivity analysis.
Overview
Using the decomposition from the event-study note, with : the pre-test examines , while the assumption actually needed is . Roth et al. (2023, §4.4) list four problems with this practice:
- Parallel pre-trends do not imply parallel post-trends. Kahn-Lang & Lang’s example: boys’ and girls’ average heights evolve in parallel until about 13 and then diverge — not a causal effect of bar mitzvahs.
- Low power. Failing to reject is not evidence of absence. As Bilinski & Hatfield put it, pre-tests “reverse the traditional roles of type I and type II error”: parallel trends is the null, so the 5% guarantee protects against falsely finding a violation, not against missing one.
- Pre-test bias. Conditioning the analysis on passing distorts estimation and inference.
- No guidance after rejection. “With enough precision, we will nearly always reject that the parallel trends assumption holds exactly,” yet a small violation may still allow useful inference.
Roth (2022) quantifies (2) and (3).
Main Content
Roth's finite-sample normal model and the NIS pre-test ^def-pretest-model
with partition . The most common test in practice passes when no individual lead is significant:
In Roth’s survey all 12 papers show pointwise CIs, 5 of 12 explicitly discuss individual significance, only one reports a joint test, and none discusses what magnitude of pre-trend the data could rule out. Normality is imposed exactly so that any distortion is attributable to trends or pre-testing rather than to asymptotic approximation; it covers TWFE, Callaway–Sant’Anna, Sun–Abraham and other asymptotically normal event-study estimators (Remark 1).
Pitfall 1 — low power
For each paper Roth finds the linear-trend slope () at which the NIS test rejects with probability 0.5 or 0.8 (), then computes the bias and coverage of the usual estimator with CI .
- Under — the conventional “minimum detectable” benchmark — the bias in the average post-period effect “is often of a magnitude comparable to, and in some cases larger than, the estimated treatment effect” (Fig. 1).
- Null rejection rates of nominal 5% tests for (Table 2, unconditional): under they range from 0.09 to 0.76; under from 0.14 to 0.98. In the most extreme case a 95% CI covers the truth only 24% of the time.
Two-period intuition ^ex-pretest-symmetry
One lead, one lag (), equal variances, true effect zero, linear trend so . By symmetry, the probability that the lead’s CI excludes zero equals the probability that the lag’s CI excludes zero. So a trend the pre-test catches half the time also generates a spuriously significant “treatment effect” half the time — “ten times more often than a 95 percent CI is supposed to reject.” More post-periods make it worse (linear bias grows with horizon); more pre-periods help — but only if early pre-periods are informative, which they are not when “treatment status [is] determined only by events close to the time of treatment.”
Non-linear violations can be worse still: a convex (e.g. exponential) differential trend is small pre-treatment, hence rarely detected, but large post-treatment; a concave trend is the benign case.
Pitfall 2 — pre-test bias
Proposition 1 — conditional mean after pre-testing ^thm-pretest-prop1
For any acceptance region ,
Proof sketch. is uncorrelated with, hence (by normality) independent of, ; take conditional expectations of .
Three terms: the target, the unconditional trend bias, and a pre-test bias equal to the regression of post on pre coefficients times the truncation-induced shift in the mean of the leads. Corollary: if parallel trends truly holds () and the test is symmetric, remains unbiased after pre-testing.
Proposition 2 — bias is exacerbated under monotone trends ^thm-pretest-prop2
Assumption 1: has common diagonal and common off-diagonal , — implied by homoskedastic errors in the non-staggered TWFE event study, where gives . If elementwise and (an upward differential trend), then
(and symmetrically for downward trends). Intuition: with a true negative lead, draws that pass the test have leads that are unusually high (close to zero); since leads and lags are positively correlated through the shared reference period, the lags in those draws are unusually high too.
Propositions 3-4 — variance shrinks after pre-testing ^thm-pretest-prop34
and if is convex (true for individual and joint significance tests), the conditional variance is weakly smaller. So under true parallel trends conventional CIs tend to over-cover after a pre-test; under violated parallel trends the bias dominates and they under-cover.
Empirically (Table 3), the additional bias from conditioning, as a percentage of the unconditional bias under , ranges from to for the first-period effect and from to for the average effect ; it has the same sign as the trend bias in all but two () or three () of the twelve papers. The share is larger for early periods because trend bias grows with horizon while pre-test bias need not.
Pre-tests as a publication filter
With a fraction of studies having violation and the rest none (eq. 4):
Screening helps only if the first factor is small enough. It converges to one — screening is useless or harmful — when (low ex-ante credibility) or when the Bayes factor (an underpowered test). This is a selection mechanism in the same family as publication bias and the “garden of forking paths.”
What to do instead (Roth 2022 §III; Roth et al. 2023 §4.4.1, §4.6)
- Report power. Use the
pretrendsR package to compute the slope (or a hypothesised non-linear path) detectable with 50%/80% power and the implied bias — a design-stage exercise akin to Power Analysis and Sample Size. - Non-inferiority / equivalence pre-tests (Bilinski & Hatfield; Dette & Schumann): test so that a large pre-trend is detected with probability . Better, but still no guarantee for the treatment-effect CI and still a pre-test.
- Avoid the pre-test altogether: Freyaldenhoven, Hansen & Shapiro’s covariate-proxy approach, or Rambachan & Roth’s honest confidence sets.
- Construct rather than test parallel trends: SDID reweights units and periods so that trends are parallel by design and still delivers valid large-panel inference; its authors note that it “addresses pretesting concerns recently expressed in Roth [2018]” (Arkhangelsky et al. §1) — see SDID vs DiD vs Synthetic Control.
- Bring context. “Bringing economic knowledge to bear on how parallel trends might plausibly be violated … will yield stronger, more credible inferences than relying on the statistical significance of pre-trends tests alone.”
Examples
Selected rows from Roth’s survey (Table 1: observed leads; Table 2: null rejection probability of a nominal 5% test for under ):
| Paper | # pre-periods | Max | Unconditional | Conditional on passing |
|---|---|---|---|---|
| Bailey & Goodman-Bacon (2015) | 5 | 1.67 | 0.94 | 0.95 |
| Deryugina (2017) | 4 | 1.09 | 0.84 | 1.00 |
| Deschenes et al. (2017) | 5 | 2.24 | 0.14 | 0.25 |
| Lafortune et al. (2017) | 5 | 1.38 | 0.98 | 0.99 |
| Bosch & Campos-Vazquez (2014) | 11 | 2.36 | 0.86 | 0.61 |
Caveats Roth states: the sample is published papers (selected on passing), and the calibration uses linear violations.
Simulating pre-test bias for a geo/store event study:
import numpy as np
rng = np.random.default_rng(0)
K, M, sigma2 = 4, 4, 1.0
Sigma = np.full((K + M, K + M), sigma2 / 2) + np.eye(K + M) * sigma2 / 2 # Assumption 1, rho = sigma^2/2
t = np.r_[-np.arange(K, 0, -1), np.arange(1, M + 1)] # relative time
gamma = 0.35 # differential slope, true tau = 0
b = rng.multivariate_normal(gamma * t, Sigma, size=200_000)
passed = (np.abs(b[:, :K]) <= 1.96 * np.sqrt(sigma2)).all(axis=1)
tau_bar = b[:, K:].mean(axis=1)
print("power of pre-test:", 1 - passed.mean())
print("unconditional bias:", tau_bar.mean(), " conditional on passing:", tau_bar[passed].mean())
# -> power ~0.41; unconditional bias ~0.87; conditional-on-passing bias ~1.21
# conditional bias exceeds unconditional bias (+38%), as Proposition 2 predictsMarketing reading. “The pre-period lift chart is flat, so the test is clean” is this exact pre-test. With 4–8 noisy pretest weeks, the detectable differential slope is large; over a multi-week test window a slope half that size accumulates into a lift-sized bias. Report what trend the pretest could have detected, and prefer designs (randomised geos, SDID-style reweighting) that make the question moot.
Connections
- Event Study Designs and Dynamic Treatment Effects — supplies , and .
- Differences-in-Differences — where “include leads to test for pre-trends” is first recommended.
- Identifying Assumptions for Staggered DiD — what the pre-test is meant to probe; Callaway–Sant’Anna placebo s are subject to the same critique (Roth 2022, Remark 1).
- Multiple Testing Corrections — the NIS criterion is a union of individual tests; its size and power depend on and on .
- Covariate Balance and Matching Diagnostics — the same “absence of evidence” problem arises with balance tests after matching.
- Sensitivity Analysis in Observational Studies — the general alternative to assumption pre-testing.