← Latest papers
📈 economics

Valid Inference when Testing Violations of Parallel Trends for Difference-in-Differences

This paper proposes simple preliminary tests and corresponding confidence intervals that ensure valid statistical inference for Difference-in-Differences estimates by overcoming the bias and coverage issues of existing methods, relying on a "conditional extrapolation assumption" to link pre-treatment trend violations to post-treatment violations.

Original authors: Jonas M. Mikhaeil, Christopher Harshaw

Published 2026-05-12
📖 6 min read🧠 Deep dive

Original authors: Jonas M. Mikhaeil, Christopher Harshaw

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Picture: The "Parallel Trends" Problem

Imagine you are a detective trying to figure out if a new medicine (the treatment) cures a headache. You have two groups of people:

  1. The Treated Group: They take the medicine.
  2. The Control Group: They take a sugar pill.

To know if the medicine worked, you need to know what would have happened to the treated group if they hadn't taken the medicine. Since you can't see the future or the "alternate universe," you make a guess: The Parallel Trends Assumption.

The Assumption: You assume that, without the medicine, the treated group's headaches would have improved (or worsened) at the exact same rate as the control group's headaches. If their lines on a graph were running parallel before the medicine, you assume they would have stayed parallel after.

The Problem: The "Pre-Test" Trap

In the real world, researchers often look at the data before the medicine was given to see if the two groups were actually running parallel.

  • If the lines look parallel: "Great! The assumption holds. Let's calculate the effect."
  • If the lines look wobbly: "Uh oh. The assumption might be broken. Maybe we shouldn't trust the results."

The Catch: A recent wave of research has shown that doing this "pre-check" and then deciding whether to report your results creates a statistical mess. It's like a referee who lets a player keep playing only if they look like they aren't cheating, but then uses a broken stopwatch to time the race. The final results become biased, the confidence intervals (the margin of error) become too small, and you might think you found a cure when you didn't.

The Solution: A New Rulebook

The authors of this paper propose a new way to handle this "pre-check" so that the math works out correctly. They introduce a concept called the Conditional Extrapolation Assumption.

The Analogy: The Weather Forecaster

Imagine you are a weather forecaster trying to predict if it will rain tomorrow (the post-treatment period). You look at the weather today and yesterday (the pre-treatment period).

  • Old Way (Classical DID): You assume if it wasn't raining yesterday, it won't rain tomorrow. You don't check if the wind is blowing weirdly.
  • The "Bad" Pre-Test Way: You check the wind. If it looks calm, you predict rain. If it looks stormy, you stop. But your prediction tool breaks if you only use it on calm days.
  • The Authors' New Way: They say, "We have a rule: If the wind today is calm enough (below a specific threshold), we assume the wind tomorrow won't be worse than it is today."

This is the Conditional Extrapolation Assumption. It doesn't promise the wind will be calm tomorrow; it just promises that if today is calm enough to pass a test, then tomorrow won't be a hurricane.

How It Works in Practice

1. The "Severity Test" (The Gatekeeper)
Before doing any math, the researcher runs a simple test to measure how "wobbly" the lines were before the treatment.

  • They calculate a "severity score" of the wobble.
  • They compare this score to a pre-set Threshold (let's call it the "Acceptable Wobble Limit").
  • If the score is below the limit: The test passes. The researcher is allowed to proceed.
  • If the score is above the limit: The test fails. The researcher must stop and say, "We cannot trust this method because the groups were too different to begin with."

2. The "Wobbly" Confidence Interval
If the test passes, the researcher calculates the effect of the treatment. But here is the clever part: they don't just draw a tiny, precise line around their answer.

  • They draw a wider confidence interval (a bigger margin of error).
  • Why? Because they know there was some wobble, even if it was small. They add a "safety buffer" to the math to account for the possibility that the wobble might get slightly worse tomorrow.
  • If the pre-treatment wobble was tiny, the interval is narrow (like a standard test).
  • If the pre-treatment wobble was noticeable (but still passed the test), the interval gets wider to be safe.

Why This Is Better

The paper argues that this method fixes the "broken stopwatch" problem.

  • It admits uncertainty: Instead of pretending the pre-treatment data is perfect, it measures the imperfection and adds it to the final error margin.
  • It prevents false confidence: If the pre-treatment groups were too different, the test stops you from making a claim at all, rather than letting you make a claim with a falsely narrow margin of error.
  • It formalizes intuition: It turns the "gut feeling" researchers have when they look at a graph and say, "This looks okay, but not perfect," into a rigorous mathematical rule.

Real-World Examples Used in the Paper

The authors tested this on two real scenarios:

  1. Vietnam Public Services: They looked at whether recentralizing government services improved things like schools and water.
    • Result: For schools and water, the "wobble" was small enough to pass the test, so they reported results with wider, safer confidence intervals. For agricultural centers, the "wobble" was too huge, so they correctly refused to make any claim.
  2. Right-to-Carry Laws in Virginia: They looked at whether letting people carry guns changed crime rates.
    • Result: When looking at "assault" rates, the groups were moving in very different directions before the law changed. The test failed, and they stopped. When looking at "murder" rates, the groups were closer, so they proceeded with their adjusted, wider intervals.

The Bottom Line

This paper provides a new toolkit for researchers using the "Difference-in-Differences" method. It says: "Don't just check if your groups look parallel and then ignore the check. Instead, measure exactly how much they deviate, set a hard limit on what is acceptable, and if they pass, report your results with a bigger safety margin to account for that deviation."

This ensures that when a study says "We found an effect," we can trust that the math behind it hasn't been broken by the very act of checking the data first.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →