On the use of auxiliary variables in multiple imputation when estimating the average causal effect with missing data
This paper investigates the role of auxiliary variables in multiple imputation for estimating average causal effects with missing data, demonstrating through simulations and theoretical analysis that distinguishing between mediator and non-mediator variables and using compatible, flexible imputation models are essential for ensuring the recoverability and unbiasedness of causal estimates.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to solve a mystery: Does smoking cannabis as a teenager cause mental health struggles later in life?
To solve this, you need to look at a group of people, compare those who smoked to those who didn't, and see the difference in their mental health scores. This is what scientists call estimating the "Average Causal Effect."
But here's the problem: The data is messy. Some people forgot to fill out parts of their survey. Some dropped out of the study. Some answers are just missing. This is the "missing data" problem.
When data is missing, you can't just ignore the gaps. If you throw away the incomplete surveys (a method called "Complete Case Analysis"), you might end up with a biased picture, like trying to judge the whole class's test scores by only looking at the students who showed up to school that day.
The Solution: "Multiple Imputation" (The Fill-in-the-Blanks Game)
Scientists use a technique called Multiple Imputation (MI). Think of this as a smart guessing game. Instead of throwing away the missing surveys, the computer creates several "fake" versions of the missing data based on patterns it sees in the real data. It then solves the mystery on all these fake versions and averages the results.
But to guess correctly, the computer needs clues. This is where Auxiliary Variables come in.
The Clues: Auxiliary Variables
Imagine you are trying to guess a student's missing math test score.
- The Obvious Clue: You look at their other math grades.
- The "Auxiliary" Clue: You also look at how many hours they spent studying, or if they had a headache that day. These aren't the math grade itself, but they are related to it. In the paper, these are called Auxiliary Variables.
The paper investigates a very specific, tricky question: What happens if one of these "clues" is actually part of the cause-and-effect chain?
The authors split these clues into two types:
- The "Side-Clue" (Non-Mediator): A variable that helps explain the missing data but isn't caused by the thing you are studying. (e.g., A student's headache causing them to miss the test, but the cannabis use didn't cause the headache).
- The "Middle-Man" (Mediator): A variable that is caused by the thing you are studying and then causes the outcome. (e.g., Cannabis use Sleep Problems Mental Health Issues). Here, "Sleep Problems" is the missing data clue, but it sits right in the middle of the causal chain.
The Experiment: Testing the Guessing Game
The researchers ran a massive simulation (a computer experiment) to see which guessing strategies worked best under different "Missingness Maps" (diagrams showing why data went missing).
They tested three main strategies:
- The "Throw Away" Strategy (CCA): Just ignore the missing data.
- The "Add the Clue to the Final Equation" Strategy (A-CCA): Use the clues to fix the missing data by adding them directly to the final math formula.
- The "Smart Imputation" Strategy (MI with Auxiliary Variables): Feed the clues into the computer's "fill-in-the-blanks" engine before doing the final math.
The Big Findings
Here is what they discovered, translated into plain English:
1. The "Middle-Man" Trap
If your clue is a "Middle-Man" (a mediator), you cannot simply add it to your final math formula (Strategy 2).
- Analogy: If you want to know how much smoking causes lung cancer, and you add "tar in the lungs" to your formula, you are accidentally blocking the path of the smoking. You are measuring the effect of the tar, not the smoking.
- Result: This method (A-CCA) produced huge errors when the clue was a mediator. It made the detective think the smoking had no effect at all, or the wrong effect.
2. The "Smart Imputation" Wins (But be careful how you do it)
The best way to use these clues is to feed them into the "fill-in-the-blanks" engine (Multiple Imputation).
- The Flexible Approach (CART): The researchers found that using a non-parametric method (like a decision tree, or "CART") was very robust. It's like a detective who doesn't rely on rigid rules but looks at the whole picture. This method worked well even when the clues were tricky "Middle-Men."
- The Compatible Approach (A-SMCFCS): Another method worked well, provided the computer's "fill-in-the-blanks" rules were perfectly compatible with the final math rules.
3. The Danger of "Incompatible" Guessing
If you use a standard guessing method (like linear regression) and you don't make sure the rules for guessing match the rules for the final math, you get biased results. It's like using a ruler to measure weight. The paper shows that simply throwing in a "Middle-Man" clue into a standard guessing engine without special care can break the whole investigation.
The Takeaway for the Detective
If you are trying to figure out cause-and-effect with messy data:
- Don't just throw away missing data.
- Don't just throw every clue you have into your final math equation. If a clue is a "Middle-Man" (caused by the exposure), putting it in the final equation will ruin your results.
- Use a "Smart Imputation" tool. Feed your clues (both side-clues and middle-men) into the computer's guessing engine before you do the final math.
- Use flexible tools. The paper suggests that flexible, non-parametric tools (like CART) or specially designed "compatible" tools (A-SMCFCS) are the safest bets to avoid getting the wrong answer.
In short: To solve the mystery of missing data, you need the right clues, but you must feed them to the computer in the right way, or you'll end up solving the wrong case.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.