← Latest papers
📊 statistics

A Comparative Study of Methods for Handling Missing Data in Longitudinal Data with Implications for Causal Inference

This study utilizes Monte Carlo simulations and empirical validation to identify optimal combinations of missing data imputation and confounder control strategies for longitudinal causal inference, demonstrating that multiple imputation or random forest paired with doubly robust estimation yields the best performance for heterogeneous effects without unmeasured confounding.

Original authors: Yemian Li, Yuhui Yang, Weiwei Hu, Zonghao Li, Fangyao Chen

Published 2026-06-25
📖 6 min read🧠 Deep dive

Original authors: Yemian Li, Yuhui Yang, Weiwei Hu, Zonghao Li, Fangyao Chen

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to figure out if drinking a glass of wine every day actually makes people happier or sadder. You decide to track a group of people over several years to find the answer. This is what researchers call a "longitudinal study."

However, life gets messy. People forget to fill out surveys, they move away, or they just stop answering questions about their mood. This is missing data. At the same time, there are other things influencing the results, like how much money they make or how old they are. These are confounders (the "hidden variables" that muddy the waters).

This paper is like a giant, high-tech cooking competition. The researchers didn't just cook one dish; they simulated thousands of different scenarios to see which "recipe" (statistical method) works best when ingredients go missing and the kitchen is chaotic.

Here is the breakdown of their findings in simple terms:

The Problem: Two Messes at Once

In real-world studies, you usually have two problems happening at the same time:

  1. Missing Data: Some people dropped out or skipped questions.
  2. Confounding: It's hard to tell if the alcohol caused the mood change, or if something else (like stress or income) caused both.

Most previous studies tried to fix these problems one at a time. This study asked: "If we have a messy dataset, what is the absolute best combination of tools to fix the missing pieces AND separate the real cause from the noise?"

The Experiment: The Simulation Kitchen

The researchers built a virtual world using a computer. They created fake people with fake lives, fake drinking habits, and fake moods. Then, they deliberately "broke" the data by:

  • Making some people disappear (missing data).
  • Making the disappearances happen for different reasons (some random, some because they were unhappy).
  • Changing the number of years they were tracked.
  • Adding different numbers of "hidden variables" (confounders).

They then tested 7 different ways to fix the missing data (like guessing the missing numbers) and 6 different ways to control for the hidden variables (like adjusting the recipe). They ran this simulation 1,000 times for each combination to see which one gave the most accurate answer.

The Ingredients: The Methods Tested

To fix missing data (The "Fill-in-the-Blanks" team):

  • Complete Case Analysis: Just throw away anyone with a missing piece. (Like throwing away a whole cake because one egg is missing).
  • Mean Imputation: Fill the gap with the average. (Like guessing a missing test score is the class average).
  • Multiple Imputation (MI): Create several different "what-if" versions of the missing data and average the results. (Like asking 10 different chefs to guess the missing ingredient, then taking the average guess).
  • Random Forest & KNN: Smart computer algorithms that look at similar people to guess the missing values.

To control for confounders (The "Noise-Canceling" team):

  • Multivariate Adjustment (MA): Simply adding all the known variables into the math equation.
  • Propensity Score Methods (PSM, IPW): Trying to match drinkers with non-drinkers who look exactly the same on paper, or giving more weight to rare types of people.
  • Doubly Robust Estimation (DRE): A "double-check" system. It uses two different math models; if one is wrong, the other saves the day.
  • Instrumental Variable (IV): Using a third factor (like a genetic trait or a random event) that influences drinking but not mood directly, to act as a natural experiment.

The Results: Who Won the Competition?

1. The "Missing Mechanism" Myth
The researchers found that why the data was missing (whether it was random or because people were sad) didn't actually change the final answer much. The most important thing was how much data was missing. The more missing data, the worse the results, no matter which method you used.

2. The Best "Fill-in-the-Blanks" Team
For causal inference (finding the true cause), the winners were:

  • Multiple Imputation (MI): The gold standard. It handles uncertainty well.
  • Mean Imputation: Surprisingly good for this specific goal, even though it's simple.
  • Random Forest: The smart algorithm that learns from patterns.

Avoid: Simply throwing away people with missing data (Complete Case Analysis) was the worst performer. It created the most errors.

3. The Best "Noise-Canceling" Team
This depended on the situation:

  • Scenario A (Complex, Real-World Effects): If the effect of alcohol varies from person to person (heterogeneous) and there are no hidden secrets, Doubly Robust Estimation (DRE) was the clear winner. It was the most accurate and stable.
  • Scenario B (Simple Effects or Hidden Secrets): If the effect is the same for everyone, or if there are unmeasured confounders (secrets the researchers didn't know about), the Instrumental Variable (IV) method was the best. It's the only one that could peek behind the curtain of hidden variables.

4. The Golden Combination
The study found that the best strategy for most real-world scenarios (where effects vary and we assume we know the main confounders) is:

Multiple Imputation (or Random Forest) + Doubly Robust Estimation.

This combination was like having a master chef (MI) and a double-check system (DRE). It produced the most accurate results with the least amount of error.

The Real-World Taste Test

To prove their simulation wasn't just a computer fantasy, they applied these methods to real data from the China Health and Retirement Longitudinal Study (CHARLS). They looked at the link between alcohol and depression in older Chinese adults.

The result? The real-world data behaved exactly like their simulation.

  • All methods found a link between alcohol and depression.
  • The "Doubly Robust" methods combined with "Multiple Imputation" gave the most stable and reliable numbers.
  • The "Throw away missing data" method (Complete Case) combined with "Inverse Probability Weighting" gave the most erratic and inflated results.

The Bottom Line

If you are a researcher trying to figure out cause-and-effect in long-term studies:

  1. Don't just delete missing data. Use smart filling methods like Multiple Imputation or Random Forest.
  2. Use a "Double-Check" system. If you think the effect varies between people, use Doubly Robust Estimation.
  3. Watch out for secrets. If you suspect there are hidden factors you can't measure, you might need an Instrumental Variable, but be careful—it loses power if you have too many follow-up years.

The paper concludes that by matching the right "fill-in" tool with the right "noise-canceling" tool, researchers can get much clearer, more trustworthy answers from messy, real-world data.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →