Estimation beyond Missing (Completely) at Random
This paper develops a robust framework for estimating population parameters under non-random missingness by introducing a missing-data analogue of Huber’s -contamination model and demonstrating that "realisable" contamination classes—where missingness is viewed as a biased version of a base distribution—allow for significantly improved minimax performance compared to arbitrary contamination models.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to figure out the average height of people in a massive, crowded city. To do this, you can’t measure everyone, so you take a sample.
In a perfect world, your sample is "Missing Completely at Random" (MCAR). This is like if you stood on a street corner and picked people at random; the fact that some people were too busy to stop doesn't tell you anything about their height. Your sample is a tiny, perfect mirror of the city.
But in the real world, data is often "Missing Not at Random" (MNAR). This is like if you only surveyed people at a basketball court. You’ll notice your "average" height is much higher than the true city average. The "missingness" is biased.
This paper, written by a team of elite statisticians, provides a new mathematical toolkit to solve this problem. Here is the breakdown of their work using everyday analogies.
1. The "Bad Apple" Model (Arbitrary Contamination)
The researchers first look at the most pessimistic scenario. Imagine you are making a fruit salad. You expect a bowl of fresh strawberries (your target population), but someone has secretly tossed in a handful of rotten, bruised berries (the "contamination").
In statistics, this is called Arbitrary -contamination. You know a certain percentage () of your data is "garbage" or biased, but you have no idea what that garbage looks like. It could be anything.
The Paper's Contribution: They created a "Robust Mean Estimator." It’s like a high-tech fruit sorter that can look at a bowl of mixed berries and, even if it doesn't know exactly what a "bad" berry looks like, it can mathematically ignore the outliers to give you a very accurate estimate of the healthy strawberries.
2. The "Family Resemblance" Model (Realisable Contamination)
This is where the paper gets really clever. The authors argue that in most real-world cases, the "bad" data isn't just random garbage; it actually comes from the same "family" as the good data.
The Analogy: Imagine you are studying the wealth of a neighborhood. You miss data from the extremely wealthy (they don't answer surveys) and the extremely poor (they don't have phones). The "missing" people aren't random aliens; they are still humans from that same neighborhood, just at the extreme ends of the spectrum.
This is what they call "Realisable" contamination. Because the biased data follows the same general "shape" (the same base distribution) as the good data, the math becomes much more powerful.
The Breakthrough: They proved that if you assume this "family resemblance," you can achieve much higher accuracy. In fact, they showed that even if almost all your data is biased (as long as it follows that "family" rule), you can still mathematically "zoom in" and find the true average.
3. The "Smart Filter" (The Kolmogorov Estimator)
How do you actually do this? They introduced a method called the Minimum Kolmogorov Distance Estimator.
The Analogy: Imagine you have a pile of puzzle pieces. You don't know what the final picture is, but you have a set of "possible" pictures (the realisable models). Instead of just averaging the pieces, you try to find the picture that "fits" the pattern of the pieces you actually have most closely.
By mathematically "projecting" your messy, incomplete data onto the set of "possible real" patterns, you can filter out the bias and find the truth.
4. Why does this matter? (The "So What?")
This isn't just math for math's sake. This research has massive implications for:
- Medicine: If a clinical trial only records data from patients who are healthy enough to show up to the clinic, the results are biased. This math helps doctors correct that bias.
- Economics: If income surveys miss the very rich and the very poor, this math helps economists calculate a more honest "average income."
- Epidemiology: If people with certain symptoms are less likely to report them, this helps scientists understand the true spread of a disease.
Summary in one sentence:
The paper provides a mathematical "lens" that allows scientists to look at biased, incomplete, and messy data and see the clear, true picture underneath.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.