← Latest papers
📈 economics

Conformal Inference for Counterfactuals and Individual Treatment Effects with Experiment Attrition

This paper introduces a novel method that combines conformal inference with established missing data techniques to generate robust and precise prediction intervals for individual treatment effects, effectively addressing the challenges of experiment attrition where traditional approaches often fail due to strong assumptions.

Original authors: Xiangyu Song

Published 2026-04-02
📖 5 min read🧠 Deep dive

Original authors: Xiangyu Song

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a teacher trying to figure out if a new study method actually helps students get better grades. You run an experiment: half the class uses the new method (Treatment), and the other half sticks to the old one (Control).

But here's the catch: Attrition.

Some students drop out of the study halfway through. Maybe they got sick, moved away, or just got bored. Now, you have a problem. You can see the grades of the students who stayed, but you have no idea how the students who left would have performed.

The Old Way: Guessing and Weighting

Traditionally, researchers have tried to fix this in a few ways:

  1. The "Ignore Them" Approach: You only look at the students who stayed. But what if the students who dropped out were the ones struggling the most? Your results would be biased.
  2. The "Fill in the Blanks" Approach (Imputation): You try to guess the missing grades based on the students who stayed. This works well if your guess is perfect, but if your assumptions about how students learn are wrong, your whole conclusion falls apart.
  3. The "Weighting" Approach: You give more importance to the students who stayed who look like the ones who left. But this gets messy and can still be inaccurate if the "dropouts" were fundamentally different in ways you can't see.

The New Way: The "Safety Net" (Conformal Inference)

This paper introduces a new tool called Conformal Inference. Think of it not as trying to guess the exact missing grade, but as building a safety net around your predictions.

Instead of saying, "Student A would have gotten a B," the new method says, "We are 95% confident Student A would have gotten a grade between a B- and a B+."

Here is how the author, Xiangyu Song, improves this safety net for the specific problem of people dropping out of experiments:

1. The "Two-Step" Safety Net

The paper realizes there are two different kinds of missing information:

  • Step 1: For the students who stayed, we are missing one of their two potential grades (e.g., we know their grade with the new method, but not what they would have gotten with the old one).
  • Step 2: For the students who dropped out, we are missing both grades.

The author's method builds a safety net for Step 1 first, and then uses that to build a bigger, stronger safety net for Step 2.

2. The "Translator" (Handling the Shift)

Here is the tricky part: The students who dropped out might be different from the ones who stayed. Maybe the dropouts were more stressed or had less time. In statistics, this is called Covariate Shift. It's like trying to predict the weather in a desert using data from a rainforest. If you don't adjust for the difference, your prediction will be wrong.

The paper uses a clever mathematical "translator" (called a semiparametric efficient estimator) to adjust the safety net. It says, "Okay, the dropouts look different from the stayers, so let's widen or shift our safety net slightly to account for that difference." This ensures the net is still strong enough to catch the truth, even if the groups aren't identical.

3. Why It's Better

The author tested this method with computer simulations (like running the experiment a thousand times in a computer) and real-world data from political science studies.

  • The Old Methods: Either the safety net was too wide (useless because it covered everything from an F to an A+) or it was too narrow and missed the truth (the student actually got an F, but the net said they got a B).
  • The New Method: It found the "Goldilocks" zone. The safety nets were narrow enough to be useful (giving a precise range) but wide enough to be correct (catching the true result 95% of the time).

The Real-World Impact

The author re-analyzed two real studies:

  1. A study on financial markets: They found that the students who dropped out actually reacted differently to the treatment than those who stayed. If you only looked at the students who stayed, you would have missed a huge part of the story.
  2. A study on civic education in Tunisia: Similar story. The dropouts had different reactions. By using this new method, the researchers could see that the treatment effect was actually stronger (or weaker) when you included the people who left.

The Bottom Line

This paper gives researchers a new, robust way to handle the inevitable problem of people dropping out of studies. Instead of throwing away data or making wild guesses, they can now build a statistical safety net that is both precise and reliable.

It's like upgrading from a fishing rod that only catches the fish you can see, to a net that can accurately estimate the size and location of the fish that swam away, ensuring you don't miss the whole picture.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →