Pretest With Caution
Difference-in-differences rests on one assumption you can never check: that the treated and untreated groups would have moved in parallel absent treatment. Because that counterfactual is unobservable, a ritual grew up around it. Estimate the event study, look at the pre-treatment coefficients, confirm they are small and jointly insignificant, and declare parallel trends "tested and passed." Jonathan Roth's Pretest with Caution takes that ritual apart. The tests we run have almost no power against exactly the confounding trends that would bias us most, and — more subtly — the act of conditioning on a passing test warps the very sampling distribution we then use for inference. Passing is weak evidence, and it is self-selected weak evidence.
The Pre-Test Ritual
Fix the standard dynamic specification. With treatment turning on at event time zero, we estimate a set of relative-time coefficients around it:
$$ Y_{it} \;=\; \alpha_i + \lambda_t + \sum_{k \neq -1} \beta_k \, D_{it}^{(k)} + \varepsilon_{it}, $$where \( D_{it}^{(k)} \) indicates that unit \( i \) is \( k \) periods from its treatment date at time \( t \), and the period just before treatment (\( k = -1 \)) is the omitted baseline. The coefficients split into lags \( \beta_0, \beta_1, \dots \) (the dynamic treatment effects we care about) and leads \( \beta_{-2}, \beta_{-3}, \dots \) (the pre-treatment "placebo" coefficients). Under the identifying assumption, the leads should all be zero in expectation:
$$ H_0:\; \beta_{-k} = 0 \quad \text{for all } k \geq 2. $$The parallel-trends assumption is precisely that the untreated potential outcomes of the two groups follow paths that differ only by a fixed level — no differential drift. The leads are the observable shadow of that assumption in the pre-period, so testing \( H_0 \) with an \( F \)-test on the leads feels like testing parallel trends itself. It is not. Parallel trends is a statement about the post-treatment counterfactual, which the pre-period data are silent about. A clean pre-period is consistent with the assumption; it does not imply it. Kahn-Lang and Lang (2020) press this point on substantive grounds: parallel trends is rarely a natural default, the choice of functional form (levels versus logs) can make it hold or fail (Roth and Sant'Anna, 2023), and a passing pre-test does nothing to justify the economic story that the two groups would have converged.
Definition: A pre-trends test as a screening rule
Let \( \hat{\beta}_{\mathrm{pre}} = (\hat\beta_{-2}, \dots, \hat\beta_{-K}) \) be the estimated leads with covariance \( \hat\Sigma_{\mathrm{pre}} \). The conventional test rejects when the Wald statistic \( \hat\beta_{\mathrm{pre}}' \hat\Sigma_{\mathrm{pre}}^{-1} \hat\beta_{\mathrm{pre}} \) exceeds a \( \chi^2 \) critical value. In practice the design is kept — published, interpreted, acted on — only when the test fails to reject. That "fail to reject ⇒ proceed" step is a screening rule, and Roth's paper is about what screening does to everything downstream.
Low Power Where It Matters
The first problem is power. A hypothesis test that fails to reject the null is not evidence for the null unless it had a good chance of rejecting when the null was false. Pre-trend tests routinely do not. Roth (2022) shows that against the specific confounding trends that would most bias the treatment estimate — smooth, low-order differential trends, the linear trend being the canonical case — commonly used pre-tests have low power. The pathology is worst exactly where it hurts: a modest linear divergence between the groups can be very likely to slip past the test while generating substantial bias in the post-treatment coefficients, because a linear trend that is nearly invisible in a short, noisy pre-period keeps accumulating into the post-period.
Roth makes this concrete by calibrating simulations to a survey of recent event-study papers in leading economics journals. The finding is not reassuring: for many published designs, the pre-trend test would have had low power against a linear violation large enough to meaningfully distort the reported effect — in a number of cases well below the coin-flip mark of fifty percent. Low power is not an occasional accident of small samples; it is the typical state of the designs the field already accepts. The intuition is dimensional. The pre-test spreads its attention across all directions of pre-trend violation, but bias in the estimand loads on one particular direction — the trend that, extrapolated forward, moves the lags. A test tuned to detect "any" violation is badly tuned to detect "the one that matters."
⚠️ "The leads look flat" is a statement about your standard errors
Insignificant leads can mean the trends are genuinely parallel, or they can mean the pre-period is too short and too noisy to detect the trend that is quietly biasing you. The two are observationally identical in the point estimates and distinguishable only by power. Eyeballing a flat pre-trend plot, or reporting a large \( F \)-test \( p \)-value, tells you nothing about which world you are in unless you also report how large a trend the test could have caught.
A pre-trend the test can't see, biasing the effect it can't ignore
Treated and control groups sit apart in level and drift apart by a small differential linear trend. A genuine treatment effect of +8 switches on at period 0 (dashed line = the treated group's no-effect counterfactual). Raise the noise and the pre-period leads look flat to the eye and to the test — yet the same trend keeps accumulating past period 0 and contaminates the difference-in-differences estimate. The readout shows the pre-trend test on this sample and the power it actually had.
A trend too small for the pre-test to reliably reject still shoves the DiD estimate off the true +8. Turn up the noise and watch the power collapse: the test tends to “pass” precisely when it is least able to protect you.
Testing Distorts Inference
The second problem is deeper and less intuitive. Even setting power aside, the act of conditioning on having passed the pre-test changes the sampling distribution of the treatment estimator. Because the leads and the lags are estimated from overlapping data and are statistically correlated, selecting the samples in which \( \hat\beta_{\mathrm{pre}} \) happened to look small is not selecting a random subset of the world — it is selecting a subset correlated with the pre-period noise, and that correlation leaks into the post-period estimate.
Formally, the object we actually report is not the unconditional estimator \( \hat\beta_{\mathrm{post}} \) but the estimator conditional on passing:
$$ \mathbb{E}\!\left[\, \hat\beta_{\mathrm{post}} \;\middle|\; \text{pre-test not rejected} \,\right] \;\neq\; \mathbb{E}\!\left[\, \hat\beta_{\mathrm{post}} \,\right]. $$Truncating the joint distribution of \( (\hat\beta_{\mathrm{pre}}, \hat\beta_{\mathrm{post}}) \) to the region where the leads are small shifts and skews the marginal distribution of the lags. Roth's striking result is on the sign of the effect: conditioning on passing can exacerbate the bias rather than reduce it. The mental model many practitioners carry — "if there were a pre-trend, the test would have caught it, so passing means the residual bias is small" — has it backwards in the cases that matter. Under a real underlying trend, the samples that pass are disproportionately the ones where noise masked the pre-trend, and in exactly those samples the same trend is still pushing on the post-period coefficients. Passing filters toward the contaminated draws.
Coverage suffers in lockstep. A nominal 95% confidence interval built the usual way — ignoring that the design was retained only because it passed a screen — under-covers the true effect once you account for the selection. Roth documents this under-coverage in the calibrated simulations: intervals that advertise 95% do not deliver it after the pre-test gate. The confidence interval is answering the wrong question. It quantifies uncertainty in a hypothetical world where you would have reported the estimate regardless of the pre-test, but you would not have; you would have thrown it away had the leads looked bad.
💡 The mechanism in one sentence
Passing the pre-test is correlated with the pre-period noise, so conditioning on it induces a truncated, shifted sampling distribution for the treatment estimate — the standard errors and the naive point estimate are both computed as if that selection never happened.
The Pre-Test Estimator Lineage
None of this is new to statistics; it is new only to the event-study ritual. The failure is the classical pre-test estimator problem, studied since the 1940s. When you first test whether a coefficient belongs in a regression and then estimate the model implied by the test's verdict, the resulting estimator has a sampling distribution that is neither the restricted nor the unrestricted one — it is a data-dependent mixture, discontinuous in the truth, with bias and non-normal tails that the textbook standard errors ignore (Judge and Bock's line of work formalized this; the caution against "estimation after model selection" is a staple of the econometrics canon). Reporting a conventional confidence interval after a specification search is invalid for the same reason a pre-trend-gated interval is: the reported inference does not condition on the selection event that produced the estimate.
Seen this way, the event-study convention is a specific and especially seductive instance of a known trap. It is seductive because the pre-test looks like a diagnostic of the identifying assumption rather than a model-selection step — but operationally, "keep the design iff the leads are insignificant" is model selection, and it carries the pre-test estimator's full baggage. The remedy is also the classical one: stop letting a binary test silently prune the sample, and instead characterize the whole family of conclusions consistent with the data.
Diagnose the Power
If you are going to look at pre-trends at all, Roth's constructive advice is to report what the test could actually see. The companion pretrends R package (Roth and Sant'Anna) turns the abstract worry into two numbers a reader can interrogate. First, given the estimated event study and its covariance matrix, it computes the power of the pre-test against a hypothesized violation — a linear trend of a given slope, say. Second, and more usefully, it inverts the question: it finds the linear trend against which the test has a chosen power, for example the slope the test would detect only 50% of the time.
That "50% slope" is a communication device. In the package's worked example (built on He and Wang, 2017), the slope the pre-test would catch half the time is on the order of \( 0.05 \) per period — and the package then plots what that undetectable-half-the-time trend would do to the event study, together with the expected coefficients conditional on passing the pre-test under that trend (its meanAfterPretesting output). If that hypothesized trend is economically large relative to your effect, your flat leads mean little; the test simply could not have ruled it out. If it is implausibly large, the pre-test is genuinely informative. Either way you have replaced "the leads are insignificant" with "the test could rule out trends bigger than this, and here is the distortion a just-detectable trend would leave behind."
Definition: The 50%-power trend as a yardstick
Let \( \gamma \) index a family of hypothesized linear pre-trends (slope per period). The 50%-power slope \( \gamma_{0.5} \) solves \( \Pr(\text{reject} \mid \text{trend } \gamma_{0.5}) = 0.5 \). Report it next to your estimated effect. A pre-test that only reliably catches a trend an order of magnitude larger than the effect you are claiming is not a safeguard — it is decoration.
A textbook “clean” event study: every pre-treatment lead (grey) is individually insignificant — but look at the width of those intervals. Insignificant-and-wide is low power, not evidence of parallel trends; the lags (green) show the real, growing effect the design was built to measure.
Sensitivity, Not Testing
The more thorough answer abandons the binary test for a sensitivity analysis. Rambachan and Roth (2023), "A More Credible Approach to Parallel Trends," reframes the whole problem. Instead of asking "is parallel trends true or false," you posit that the post-treatment violation of parallel trends is restricted — bounded by what you saw in the pre-period — and report the entire set of treatment effects consistent with that restriction. The estimand becomes partially identified, and the deliverable is an interval of intervals rather than a point with a green checkmark.
Two restriction families do most of the work. The relative-magnitudes restriction says the post-period differential trend is no larger than \( M \) times the largest pre-period differential trend — you let the observed pre-trends set the scale of plausible post-trends, tuning the multiplier \( M \). The smoothness restriction instead bounds how much the slope of the differential trend can change from one period to the next, formalizing the idea that secular confounders evolve gradually rather than jumping at the treatment date. For any choice of restriction, the HonestDiD package returns valid confidence sets for the treatment effect — built from conditional, fixed-length, and hybridized confidence sets that remain honest under the partial identification.
The natural summary is a breakdown analysis: the largest restriction (the biggest \( M \), or the largest allowed slope change) under which you can still reject a null effect. Rather than "we passed the pre-test," you report "our conclusion survives as long as the post-treatment differential trend is no more than, say, twice the pre-treatment one." That is a claim a skeptical reader can argue with on its merits, and it degrades gracefully — the finding does not flip from "valid" to "invalid" at a \( p \)-value threshold; it weakens continuously as you allow larger violations.
The reframing in a line
A pre-test asks a yes/no question the data cannot answer. A sensitivity analysis asks "how wrong could the identifying assumption be before my conclusion breaks?" — a question the data can speak to, and the one a decision-maker actually needs answered.
A Family of Better Ideas
Roth's paper sits inside a broader movement away from the naive pre-test. Several strands are worth knowing. Bilinski and Hatfield (2018/2026), "Nothing to See Here?", make the equivalence-testing case sharp: the conventional test controls the probability of wrongly declaring non-parallel trends, which is the error you do not care about, while leaving the error you do care about — missing a real violation — uncontrolled. They propose a non-inferiority framing that instead tightly bounds the chance of missing a violation large enough to matter, measured on the scale of the treatment effect, and show it can be used as a screening step with little or no induced bias under common error structures. The philosophical inversion — put the burden of proof on parallelism, not on its absence — is the right one.
Freyaldenhoven, Hansen, and Shapiro (2019) attack the confounder directly: if a covariate proxies the trend that threatens parallelism, you can use it as an instrument-like control and estimate the treatment effect while explicitly modeling the pre-event dynamics, rather than testing them away. On the estimator side, the modern event-study literature — Callaway and Sant'Anna (2021) and Sun and Abraham (2021) — cleaned up the coefficients that feed these tests, so that the leads and lags are not themselves contaminated by the negative-weighting pathologies of two-way fixed effects under staggered adoption (the subject of our companion piece on staggered difference-in-differences). And the whole enterprise descends from the partial-identification tradition of Manski and Pepper — when a point-identifying assumption is untestable, report the bounds implied by a credible weaker assumption instead of asserting the strong one.
The common thread: replace a fragile point identification defended by an underpowered test with an honest range defended by a stated, arguable restriction. You give up the clean single number and the reassuring checkmark. You gain a result that means what it says.
Geo Lift Tests Inherit This
Marketing measurement runs on difference-in-differences whether or not it uses the name. A matched-market geo lift test picks control markets whose pre-period outcome trajectory tracks the treated markets, turns advertising on in the treated set, and reads incrementality off the post-period divergence — a DiD in all but branding. The identifying assumption is the same parallel-trends condition, and the "matching" step is the same pre-trend check: we accept the control set because the pre-period lined up. Everything above applies without modification.
The two failure modes translate directly. Low power: with a handful of markets and a few months of pre-period, the "the markets matched well before launch" check can easily miss a differential trend — a control region that was quietly decelerating, a treated region riding a local demand wave — large enough to swamp a real lift of a few percent. Conditioning distortion: when we search over candidate control sets and keep the one whose pre-period fits best, we are running exactly the pre-test-and-select procedure that biases the estimate and breaks the confidence interval. The tighter the pre-period match we demand, the more aggressively we are selecting on pre-period noise, and the more the reported lift interval overstates its own precision.
⚠️ A clean pre-period match is a starting point, not a certificate
Report the power of the match, not just the fact of it: how large a pre-existing test-versus-control trend would your matching window have failed to catch? Then bound the lift under a stated sensitivity restriction — "the conclusion holds as long as any residual geo-level trend is no more than \( M \) times what we saw pre-launch" — rather than treating a good pre-period fit as proof of comparability. Synthetic-control weighting and honest bounds beat a green/red pre-check.
This is why the framework leans on randomization where it can and on sensitivity where it cannot. A properly randomized geo experiment (see geo lift and time-based regression) earns parallel trends by design — random assignment makes the treated and control groups exchangeable in expectation, so there is no untestable trend to pre-test. When randomization is impossible and matched-market DiD is the only option, the honest move is to carry the identifying assumption forward explicitly: state it, bound the effect under credible violations of it, and refuse to let an underpowered pre-period check stand in for the argument. The measurement platform's continuous-learning loop treats a passed pre-trend check as one input to a sensitivity band, never as a license. In observational MMM, the same discipline shows up as calibrating the model against lift experiments and reporting how far the conclusions can bend — the difference-in-differences instinct, held honestly.
Takeaways
- Parallel trends is the untestable identifying assumption of DiD; a clean pre-period is consistent with it but does not imply it. Testing the leads is not testing the assumption.
- Pre-trend tests have low power against the smooth, low-order differential trends — linear trends above all — that would most bias the treatment estimate. Calibrated to published designs, that power is often well below 50%.
- Conditioning on passing the pre-test distorts inference: it truncates and shifts the estimator's sampling distribution, can exacerbate bias in the retained samples, and makes nominal 95% intervals under-cover. This is the classical pre-test estimator problem.
- If you look at pre-trends, report the power: use the
pretrendspackage to state the trend the test would catch only half the time and the distortion a just-undetectable trend would leave. - Prefer sensitivity analysis (Rambachan–Roth
HonestDiD): bound post-period violations by the pre-period (relative magnitudes) or by smoothness, and report the breakdown point where the conclusion fails — not a binary pass/fail. - Geo lift tests and matched-market DiD inherit all of this. A pre-period match is a starting point, not a certificate; randomize where possible, and bound the lift under stated restrictions where not.
References
- Roth, J. (2022). Pretest with Caution: Event-Study Estimates after Testing for Parallel Trends. American Economic Review: Insights, 4(3), 305–322. Working paper: jonathandroth.com.
- Rambachan, A., & Roth, J. (2023). A More Credible Approach to Parallel Trends. Review of Economic Studies, 90(5), 2555–2591.
- Roth, J., & Sant'Anna, P. H. C. pretrends (R package): power calculations for pre-trends tests. github.com/jonathandroth/pretrends.
- Rambachan, A., & Roth, J. HonestDiD (R package): robust inference in DiD and event-study designs. github.com/asheshrambachan/HonestDiD.
- Bilinski, A., & Hatfield, L. A. (2018/2026). Nothing to See Here? A Non-Inferiority Approach to Parallel Trends. Statistics in Medicine; arXiv:1805.03273.
- Freyaldenhoven, S., Hansen, C., & Shapiro, J. M. (2019). Pre-event Trends in the Panel Event-Study Design. American Economic Review, 109(9), 3307–3338.
- Kahn-Lang, A., & Lang, K. (2020). The Promise and Pitfalls of Differences-in-Differences: Reflections on 16 and Pregnant and Other Applications. Journal of Business & Economic Statistics, 38(3), 613–620.
- Roth, J., & Sant'Anna, P. H. C. (2023). When Is Parallel Trends Sensitive to Functional Form? Econometrica, 91(2), 737–747.
- Callaway, B., & Sant'Anna, P. H. C. (2021). Difference-in-Differences with Multiple Time Periods. Journal of Econometrics, 225(2), 200–230.
- Sun, L., & Abraham, S. (2021). Estimating Dynamic Treatment Effects in Event Studies with Heterogeneous Treatment Effects. Journal of Econometrics, 225(2), 175–199.
- Manski, C. F., & Pepper, J. V. (2018). How Do Right-to-Carry Laws Affect Crime Rates? Coping with Ambiguity Using Bounded-Variation Assumptions. Review of Economics and Statistics, 100(2), 232–244.