← Latest papers
📈 economics

The Generalized Falsification Adaptive Set for Violations of the Exclusion Restriction and Exogeneity

This paper extends the Falsification Adaptive Set (FAS) framework by demonstrating that the nature of instrument invalidity (confounder vs. collider) critically affects identification, and proposes a generalized FAS that guarantees inclusion of the true parameter if at least one instrument remains valid.

Original authors: Nicolas Apfel, Frank Windmeijer

Published 2026-07-02
📖 5 min read🧠 Deep dive

Original authors: Nicolas Apfel, Frank Windmeijer

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to solve a mystery: What is the true effect of a specific action (like building a highway) on an outcome (like a city's trade)?

To solve this, you can't just look at the data directly because other hidden factors might be messing up the picture. So, you use "clues" called Instruments. These are things that affect the action but shouldn't directly affect the outcome, except through that action.

However, sometimes your clues are flawed. They might be "invalid." This paper is about what happens when your clues are broken, and how to still find the truth without getting tricked.

Here is the breakdown of the paper's ideas using simple analogies:

1. The Two Ways a Clue Can Be Broken

The authors say that when a clue (instrument) is bad, it fails in one of two very different ways. Think of them as two types of "bad actors":

  • The "Meddler" (Violating Exclusion): This clue is actually a direct cause of the outcome, not just the action.
    • Analogy: Imagine you are trying to see if Rain causes Traffic. You use Clouds as your clue. But what if the clouds are actually carrying a heavy load that crushes the cars? The clouds aren't just predicting rain; they are directly crushing the cars. In this case, you must include the clouds in your analysis to control for their direct damage.
  • The "Colluder" (Violating Exogeneity): This clue is influenced by the hidden secrets (the error term) you are trying to avoid.
    • Analogy: Imagine you are trying to see if Studying causes Good Grades. You use Sitting in the Front Row as your clue. But what if the teacher only puts the smartest students (who would have gotten good grades anyway) in the front row? The clue is "colluding" with the hidden smartness. In this case, you must exclude the clue entirely, or it will trick you.

The Big Problem: The previous method (by Masten and Poirier) assumed all bad clues were "Meddlers." The authors show that if you treat a "Colluder" like a "Meddler" (or vice versa), your final answer will be wrong, and you might accidentally throw out the true answer.

2. The "Falsification Adaptive Set" (FAS)

When your model is "falsified" (meaning the clues don't fit together perfectly), you can't give a single number for the answer. Instead, you give a range of possible answers that could be true if you relax your rules just a tiny bit.

The authors created a new tool called the Generalized Falsification Adaptive Set (Generalized FAS).

  • The Old Way: You had to guess how the clues were broken. If you guessed "Meddler," you got one range. If you guessed "Colluder," you got a different range. If you guessed wrong, the true answer might be outside your range.
  • The New Way (Generalized FAS): The authors say, "Let's not guess." Instead, let's build every possible range for every possible combination of broken clues. Then, we take the Union (the big umbrella) of all those ranges.

The Result: As long as you have at least one good clue that is relevant, this big umbrella is guaranteed to catch the true answer, no matter how the other clues are broken. You don't need to know exactly which clues are "Meddlers" and which are "Colluders" beforehand.

3. The "Roads and Trade" Example

To prove this works, the authors looked at a real study about Highways and Trade.

  • They used three historical clues: 1947 highway plans, 1898 railroad routes, and old exploration paths.
  • The data showed the clues didn't fit perfectly (the model was "falsified").
  • They calculated the ranges for all possible ways these clues could be broken.
  • The Finding: The "Generalized FAS" was a wide range (from -0.61 to 3.74). It was wide because it covered all possibilities. However, it was the only set that was guaranteed to contain the true effect, even if the researchers didn't know exactly which historical route was a "Meddler" and which was a "Colluder."

4. The Takeaway

  • Don't force a square peg into a round hole: You can't treat all bad clues the same way. Some need to be controlled for; others need to be thrown out.
  • The Safety Net: If you aren't sure which clues are broken and how, don't pick just one method. Use the Generalized FAS. It's like a safety net that catches the truth even if you don't know exactly where the holes in your logic are, provided you have at least one solid clue to start with.
  • Beware of Cherry-Picking: The paper warns researchers not to just pick the narrowest range that looks nice. If you pick a narrow range based on a guess about how the clues are broken, you might miss the truth. It's better to report the wider, safer range that accounts for all possibilities.

In short: When your tools are broken, don't guess how they are broken. Build a net that covers every way they could be broken, so you are guaranteed to catch the truth.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →