← Latest papers
📈 economics

Randomization Inference For the Always-Reporter Treatment Effect

This paper proposes a finite-sample valid randomization inference procedure for estimating the average treatment effect among always-reporters in randomized controlled trials with attrition, utilizing a worst-case approach that maximizes p-values over all plausible always-reporter configurations consistent with the observed data.

Original authors: Haoge Chang, Zeyang Yu

Published 2026-03-27
📖 5 min read🧠 Deep dive

Original authors: Haoge Chang, Zeyang Yu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to solve a mystery: Did a new medicine actually work?

You run a clinical trial. You give the medicine to Group A and a placebo to Group B. But here's the catch: some people disappeared. They stopped answering the phone, moved away, or simply refused to fill out the final survey. This is called "attrition."

In the real world, people who drop out often do so for a reason. Maybe the medicine made them feel sick, so they quit. If you only look at the people who stayed, you might think the medicine is great, when in reality, it was terrible for the people who left. This is a selection bias.

The "Always-Reporter" Detective Work

The authors of this paper, Haoge Chang and Zeyang Yu, propose a clever way to solve this mystery without making dangerous guesses.

They focus on a specific group of people: the "Always-Reporters."

  • The "Always-Reporters": People who would have filled out the survey whether they got the medicine or the placebo.
  • The "If-Reporters": People who only filled out the survey if they got the medicine (maybe they felt great, or maybe they were bribed).
  • The "Never-Reporters": People who never filled out the survey, no matter what.

The Problem: You can see who filled out the survey, but you cannot see who is an "Always-Reporter" and who is an "If-Reporter" just by looking at the data. It's like trying to identify a spy in a crowd; you see them acting, but you don't know their true allegiance.

The "Worst-Case" Strategy

Most statistical methods try to guess the answer by making assumptions (e.g., "Let's assume the dropouts were random"). The authors say, "No guessing. Let's be paranoid."

They use a strategy called Randomization Inference, which is like running a simulation of the universe over and over again to see what could have happened.

Here is their creative twist: The "Worst-Case" Scenario.

Imagine you are a lawyer defending a client. You don't know exactly who the "Always-Reporters" are. So, you ask the judge:

"Your Honor, let's assume the worst possible arrangement of people. Let's assume that every single possible combination of 'Always-Reporters' that fits the data we have is true. If the medicine still looks effective even in the worst possible version of reality, then we can be 100% sure it works."

They calculate a "p-value" (a measure of doubt) for every single possible scenario of who the reporters are. Then, they pick the highest p-value (the one that makes the medicine look the least effective).

  • If even this "worst-case" p-value is low enough to reject the idea that the medicine does nothing, then the medicine definitely works.
  • This guarantees that your conclusion is valid, even if you don't know exactly who dropped out and why.

The Two-Step "Pruning" Process

Checking every single possibility is like trying to read every book in a library to find one specific sentence. It takes forever. To speed this up, the authors add a "Pre-test" (a pruning step).

Think of it like a bouncer at a club:

  1. The Bouncer (Pre-test): Before we let a "scenario" into the main club (the final calculation), we check if it makes sense. For example, if a scenario suggests that 99% of the "Always-Reporters" are in the treatment group but only 1% are in the control group, the bouncer kicks it out. Why? Because the treatment was assigned randomly; the groups should be roughly equal.
  2. The Main Club (The Calculation): We only run the heavy math on the scenarios that passed the bouncer's check.

The "Magic" Statistics

To measure if the medicine worked, they use two types of "thermometers" (statistics):

  1. The Hajek Thermometer: Measures the difference in average outcomes between the two groups.
  2. The Chi-Square Thermometer: Measures if the number of people reporting is balanced between the groups.

They combine these into a super-thermometer. If the reading is high enough, the medicine is a success.

Why This Matters (The "So What?")

  • No More "Maybe": Old methods relied on big-sample math (asymptotics) that only works when you have thousands of people and the data behaves perfectly. This new method works even with small groups and messy data.
  • Finite-Sample Guarantee: It's like having a seatbelt that works even if you crash at 5 mph, not just 100 mph. It gives you a mathematically proven guarantee that you won't be fooled by the missing data.
  • Computational Magic: They figured out how to do this complex math using "Integer Programming" (a type of puzzle-solving algorithm) so that computers can actually solve it in a reasonable amount of time, rather than taking a million years.

The Analogy Summary

Imagine you are trying to judge a cooking contest, but some judges left early.

  • Old Method: "I'll assume the judges who left were just hungry and didn't like the food. So, I'll ignore them." (Risky!)
  • This Paper's Method: "I will imagine every possible reason those judges left. I will assume the worst: maybe they hated the food so much they ran away. I will calculate the score based on that worst-case scenario. If the chef still wins even when the judges hated the food, then the chef is a genius."

This paper gives us a rigorous, "paranoid" way to trust our conclusions even when our data is incomplete, ensuring that we don't accidentally declare a failure a success (or vice versa) just because some people didn't show up.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →