Learning Who to Treat When Treatment is Missing
This paper addresses the challenge of missing treatment data in policy learning by extending efficient estimators for average and conditional treatment effects under missing at random (MAR) and missing completely conditionally at random (MCCAR) assumptions, proving that MAR-based estimators are both valid and more efficient while demonstrating through experiments that correctly specifying the missingness mechanism is crucial for achieving near-oracle performance.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are the captain of a spaceship with a limited supply of fuel, and your mission is to save as many planets as possible. You have a map showing which planets are in trouble, but there's a catch: some of your sensors are glitchy. Sometimes, the sensor tells you a planet is in trouble; other times, it just goes blank. You need to decide which planets to visit to get the most "saved" points. This is the world of policy learning, a branch of computer science where algorithms help humans make the best decisions when resources are scarce. To do this, the computer needs to know the "treatment effect"—basically, how much better a planet does if you visit it compared to if you ignore it. But what happens when your data is messy? What if the sensor didn't just fail randomly, but failed more often when the planet was already in a really bad state (or a really good one)? If you ignore the planets with the broken sensors, you might miss the ones that need help the most, leading to a mission failure. This paper tackles the tricky problem of how to make smart decisions when your data on who got "treated" is incomplete or missing.
The authors of this paper, researchers from Carnegie Mellon University, are tackling a very specific headache: what do you do when you don't know if a person received a treatment (like a medicine or a social program) because the record is missing? In many real-world situations, like healthcare or social services, data isn't perfect. Sometimes, records are lost, or people forget to fill out forms. The paper looks at two main ways this missing data can happen. The first is like a broken camera that misses pictures completely at random; the second is more sneaky: the camera misses pictures specifically when the scene is chaotic or when the subject is doing something specific.
The researchers found that the "sneaky" missingness is actually the more common and dangerous scenario. They proved mathematically that if you just throw away the data where the treatment is missing (a common habit called "complete-case analysis"), you are leaving valuable information on the table. Instead, they developed a new method that acts like a detective, using the clues from the people whose data is complete to guess what might have happened to the people whose data is missing. They showed that their method, which assumes the missingness depends on what we can see (like a person's age or income), is not only safer but also more efficient. It's like having a superpower to use every single piece of information you have, rather than tossing half your puzzle pieces in the trash just because the picture on them is smudged.
Through a series of computer simulations and tests on real-world datasets (like voting records and school test scores), the team demonstrated that their approach works. When the missing data was truly random, their method performed just as well as the old way. But when the missing data was related to the outcome (the "sneaky" kind), the old way failed, giving biased and wrong answers, while their new method stayed accurate. They even showed that their method gets better and more precise as you add more data, eventually reaching a "near-perfect" level of performance. The bottom line is that for anyone trying to decide who to help with limited resources, ignoring missing data is a bad idea. By using their new tools, policymakers can make smarter, fairer choices even when their records are imperfect.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.