← Latest papers
📊 statistics

Towards Efficient Inference under Nonmonotone Missingness with General Imputation

This paper introduces the Restricted ANOVA hierarchY (RAY) estimator, a novel functional decomposition method that provides a computable, closed-form approximation to the semiparametrically efficient estimator for parameter inference under nonmonotone missingness, while offering theoretical guarantees on efficiency and adaptability across various imputation strategies.

Original authors: Qi Xu, Lorenzo Testa, Jing Lei, Kathryn Roeder

Published 2026-07-29
📖 4 min read☕ Coffee break read

Original authors: Qi Xu, Lorenzo Testa, Jing Lei, Kathryn Roeder

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to solve a giant jigsaw puzzle, but you don't have the whole picture in front of you. In the world of data science, this is a daily reality. Scientists collect information from many different sources—like measuring a person's genes, their blood pressure, and their sleep habits all at once. But often, the data is messy. Maybe one person has their genes and blood pressure recorded but forgot to log their sleep. Another person has sleep and genes but missed the blood pressure. This is called "missing data."

When the missing pieces are scattered randomly and don't follow a simple pattern (like "everyone who missed genes also missed blood pressure"), statisticians call it "nonmonotone missingness." It's the most annoying kind of puzzle because you can't just throw away the incomplete pieces; that would waste a huge amount of information. You also can't just guess the missing numbers and pretend they are real, because bad guesses can lead to wrong conclusions. The goal is to use every single scrap of information you have, even the messy ones, to get the most accurate answer possible without tricking yourself into thinking you know more than you do.

This is exactly the problem tackled in the paper "Towards Efficient Inference under Nonmonotone Missingness with General Imputation." The authors, a team from Carnegie Mellon University and Italy, introduce a new method called RAY (Restricted ANOVA hierarchY). Think of RAY as a clever new way to organize the puzzle pieces. Instead of trying to force a perfect fit or guessing blindly, RAY breaks the problem down into a hierarchy of smaller, manageable chunks. It uses a mathematical trick to figure out how much weight to give to each different type of incomplete data pattern.

The paper shows that RAY is a "safe" method. Even if the guesses (imputations) you use to fill in the blanks are imperfect or slightly wrong, RAY still gives you a valid answer. It's like having a safety net: if your guess is good, RAY uses it to make your result super precise; if your guess is bad, RAY ignores the bad parts and falls back on the solid data you already have, so you don't get a wrong answer. The researchers proved mathematically that under certain conditions (specifically when data is missing completely at random), RAY is as good as it gets—it hits the "efficiency lower bound," meaning you can't possibly get a more precise answer with the same amount of data.

They also created an "adaptive" version called aRAY. Imagine RAY is a smart car that drives well, but aRAY is a self-driving car that learns which route is fastest in real-time. aRAY automatically figures out the best way to combine all the different data patterns to get the smallest possible error. In their tests, which included simulating thousands of scenarios and applying the method to real single-cell biology data (measuring proteins in individual cells), aRAY consistently outperformed older methods. It reduced the error in estimates and made the confidence intervals (the range where the true answer is likely to be) much tighter.

However, the paper is careful to note its limits. The method works best when the missing data is random. If the data is missing because of a specific reason related to the values themselves (like sick people being less likely to show up for a checkup), the method needs extra adjustments and might not work as perfectly. The authors show through simulations that if you ignore these specific reasons for missing data, your results can become biased. But for the vast majority of cases where data is just randomly scattered, RAY and aRAY offer a powerful, mathematically proven way to turn a messy, incomplete puzzle into a clear, high-definition picture.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →