Donut Regression Discontinuity Designs
This paper provides a theoretical framework for donut regression discontinuity designs, demonstrating how excluding observations near the cutoff impacts bias and variance while showing that existing bias-aware inference methods remain valid and formalizing specification tests for comparing conventional and donut estimators.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to figure out why something happened, but you can't run a controlled experiment. You have to rely on "observational data"—just watching what happens in the real world. In the world of economics and social science, there is a clever trick called a Regression Discontinuity (RD) design. Think of it like a strict bouncer at an exclusive club. The bouncer has a rule: "If you are 18 years old or older, you get in; if you are 17 years and 364 days old, you stay out."
Because the rule is so sharp, the people who are just barely old enough (18 years and 1 day) and the people who are just barely too young (17 years and 364 days) are practically identical twins in every way except for that one day of age. If you compare their outcomes—say, how much money they make later in life—you can isolate the effect of being 18. It's like a natural experiment where the only thing that changes is the "treatment" (getting into the club).
But here's the catch: sometimes the data gets messy. People might lie about their age, or the bouncer might get confused and let in a 17-year-old who looks 18, or maybe people who are exactly 18 years old are just weirdly different from everyone else. When researchers get suspicious that the data right at the "cutoff" line is broken or manipulated, they often use a strategy called a "Donut" RD design. This is where they throw away all the data points that are too close to the line, creating a hole in the middle of their data—like a donut. They then try to guess what would have happened at the center by looking at the people who are a bit further away.
The problem is, for a long time, nobody really knew the math behind this "donut" hole. Researchers were just guessing how big the hole should be and whether their results were trustworthy. They were eating the donut without knowing if the hole was too big or if the glaze was melting.
The Paper's Story: The Donut Math
In this paper, Claudia Noack and Christoph Rothe decide to stop guessing and start doing the heavy math. They treat the "donut" not just as a heuristic trick, but as a formal statistical tool. They ask: If we cut out the middle of the data, what happens to our estimates? Do we get more wrong answers? Do our confidence intervals (the range where we think the truth lies) get wider?
Their findings are a mix of "be careful" and "here's how to do it right."
First, they discovered that cutting out the data near the line does make your estimates less precise. It's like trying to guess the temperature at the center of a room by only measuring the corners. You have to guess (extrapolate) more, which introduces more error. They found that if you remove a small "donut" (say, observations within 10% of your usual data window), your estimate's bias (how far off you might be) can jump by 41% to 63%, and your variance (how much your answer might bounce around if you did the study again) can increase by 53% to 61%. In plain English: the donut makes your answer noisier and potentially more skewed.
However, the authors also found a silver lining. They showed that a specific type of statistical tool, called "bias-aware" confidence intervals, works perfectly fine even with the donut hole. These are special intervals that are a bit wider to account for the fact that you might be slightly off. The paper proves that you don't need to invent new math for donuts; you just need to use these existing, robust tools. But there's a cost: because of the extra uncertainty, these confidence intervals get 26% to 33% longer when you use a donut. It's like the bouncer saying, "We aren't sure exactly who gets in, so we'll give you a wider range of possibilities."
The paper also tackles the question of why researchers use donuts in the first place. Often, they use them as a "diagnostic test." If the result from the "donut" data (ignoring the messy middle) is totally different from the result using all the data, it's a red flag that something is wrong with the data near the cutoff. The authors developed new, formal tests to check this. They compared two ways of doing this check:
- The Old Way: Compare the "Donut" result to the "Full Data" result.
- The New Way: Compare the "Donut" result to a result calculated only from the data inside the hole (the "within-donut" data).
Through simulations and math, they found that the New Way is generally a better detective. It is more powerful at spotting when the data near the cutoff is actually broken. In their simulations, the new test caught the problems much more often than the old method.
Finally, the authors applied their new rules to two real-world examples. One was about teacher contracts in Pakistan, where a policy change created a cutoff. They found that when they used the donut method, the results changed drastically, suggesting the original data near the cutoff was indeed problematic. The other example was about low birthweight babies (around 1,500 grams). They confirmed that rounding errors in birth weights were messing up the data right at the 1,500g mark, and their new tests helped quantify exactly how much that mattered.
In short, Noack and Rothe didn't just say "donuts are okay." They gave researchers a manual: "If you use a donut, expect your answers to be a bit noisier and your safety ranges to be wider. But if you use these specific tools and this new testing method, you can still trust your conclusions." They turned a popular but messy kitchen trick into a precise, scientific instrument.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.