Uncertainty intervals for multilevel models with missing not at random data
This paper proposes a sensitivity analysis method for linear multilevel models with missing not at random (MNAR) data that derives bias adjustments to estimate plausible bounds for parameters of interest under weaker assumptions than missing at random, validated through simulations and an application to loneliness, physical activity, and memory trajectories.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to bake the perfect cake (your study) to understand how ingredients like loneliness and exercise affect memory over time. You have a recipe that assumes if a baker drops out of the class, they leave randomly—maybe they just forgot their apron or got a flat tire. This is called the "Missing At Random" (MAR) assumption. Most statistical tools bake the cake based on this idea.
However, in the real world, bakers often drop out for specific reasons related to the cake itself. Maybe the ones who are struggling to remember the steps (poor memory) are the ones who quit the class early. If you ignore this, your final cake will taste wrong because you only tasted the bakes from the people who stayed. This is called "Missing Not At Random" (MNAR).
This paper proposes a new way to bake the cake that admits, "We don't know exactly why people left, but we can guess how much it might have changed the taste."
Here is how the author, Minna Genbäck, breaks it down:
1. The Problem: The "Ghost" Dropouts
In long-term studies (like following people's memories over 10 years), people stop answering surveys. Standard math assumes these dropouts are random. But in health studies, sicker or frailer people often drop out more than healthy ones. If you treat them as random, your results are biased—like judging a race only by the runners who finished, ignoring those who tripped and fell because the track was too hard.
2. The Solution: The "What If" Safety Net
Instead of pretending we know exactly why people left (which is impossible to prove), the author suggests a Sensitivity Analysis. Think of this as building a "safety net" around your answer.
- The Old Way: You calculate one single number (e.g., "Exercise improves memory by 0.4 points").
- The New Way: You calculate a range. You say, "If the dropouts were slightly related to poor memory, the answer is between X and Y. If they were very related, the answer is between A and B."
This creates an Uncertainty Interval. It's not a single point; it's a zone of plausible answers that covers all the reasonable "what if" scenarios.
3. The Mechanism: The "Secret Handshake"
To do this, the author uses two models that talk to each other:
- The Memory Model: Predicts how memory changes.
- The Dropout Model: Predicts who quits the study.
In standard math, these two models are strangers; they don't talk. In this new method, the author introduces a "Sensitivity Parameter" (a secret handshake). This parameter represents the correlation between why someone's memory is changing and why they are quitting.
- If the handshake is weak (zero), it's the standard "random dropout" scenario.
- If the handshake is strong, it means people with declining memory are more likely to quit.
The author derives a mathematical formula to calculate exactly how much this "handshake" would skew the results. Then, they test a range of handshake strengths (from zero to a plausible maximum) to see how much the final answer wobbles.
4. The Real-World Test: Loneliness and Memory
The author tested this on a massive European survey (SHARE) involving over 27,000 people aged 65+. They looked at whether loneliness and physical activity affected memory scores over time.
- The Standard Result: Under the "random dropout" assumption, feeling lonely was linked to a drop in memory, and exercise was linked to a boost.
- The New Result: Even when they allowed for the possibility that people with worse memory were more likely to quit (the MNAR scenario), the results didn't change much. The "safety net" (uncertainty interval) still included the original findings.
This suggests the original conclusions were robust; they held up even when the "random dropout" assumption was relaxed.
5. The Simulation: The "Stress Test"
To prove this method works, the author ran a computer simulation. They created fake data where they knew the truth and where people dropped out specifically because their memory was bad.
- They tried to fix the bias using their new formula.
- The Result: The method worked well. It successfully corrected the "wobbly" estimates and gave them a range that captured the true answer most of the time.
The Catch (Limitations)
The author admits this method is computationally heavy. It's like trying to solve a giant 3D puzzle where the pieces keep moving. If the groups of people (clusters) are too big, the math gets too slow to run on a standard computer. Also, the method assumes the "handshake" strength is within a specific, plausible range—if you guess the range is wrong, the safety net might be too wide or too narrow.
Summary
This paper doesn't give you a magic wand to fix missing data. Instead, it gives you a ruler with a wider margin of error. It allows researchers to say, "We aren't sure exactly why people left the study, but even if we assume the worst reasonable scenario, our main conclusions about loneliness and memory still stand." It replaces a fragile, single-point guess with a sturdy, flexible range of possibilities.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.