Identification of complier and noncomplier average causal effects in the presence of latent missing-at-random (LMAR) outcomes: a unifying view and choices of assumptions
This paper provides a unifying framework for identifying complier and noncomplier average causal effects in the presence of latent missing-at-random outcomes by demonstrating that, beyond treatment assignment ignorability and latent missingness, two additional assumptions regarding the missingness mechanism and principal identification are generally required, while also proposing modifications to existing assumptions and illustrating the approach with real-world data.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to figure out if a new volunteering program actually makes people feel more fulfilled. You have a group of people who signed up (the Treatment Group) and a group who didn't (the Control Group).
But there are two big problems with your experiment:
- The "Slacker" Problem (Noncompliance): Not everyone in the Treatment Group actually did the volunteering. Some showed up and worked hard (the "Compliers"), while others signed up but did very little (the "Noncompliers"). In the Control Group, nobody was asked to volunteer, so you have no idea who would have been a hard worker if they had been asked.
- The "Ghost" Problem (Missing Data): Some people dropped out of the study or forgot to fill out their final happiness survey. Their results are missing.
This paper is about how to solve the math puzzle of figuring out the true effect of the program when you have both "Slackers" and "Ghosts."
The Two Types of People (Principal Strata)
The authors divide people into two invisible groups based on how they would behave if asked:
- Compliers: People who will do the work if asked.
- Noncompliers: People who won't do the work even if asked.
We can see who is who in the Treatment group (because they actually did it or didn't). But in the Control group, everyone looks the same because no one was asked. This makes it hard to calculate the specific effect on the "Compliers" (CACE) and the "Noncompliers" (NACE).
The "Ghost" Problem: Why Missing Data is Tricky
Usually, statisticians assume that missing data is random (like a coin flip). But here, the authors assume something more complex called LMAR (Latent Missing at Random).
Think of it like this:
- In the Treatment Group, missing data might be random.
- In the Control Group, missing data might depend on whether you would have been a Complier or a Noncomplier.
- Example: Maybe the "Compliers" (who are naturally more responsible) are more likely to fill out the survey, while the "Noncompliers" (who are less engaged) are more likely to lose the survey.
Because we can't see who is a Complier in the Control group, we can't tell if the missing surveys are skewing our results.
The Big Breakthrough: The "Two-Recipe" Puzzle
The authors realized that trying to solve this problem is like trying to bake a cake when you don't know two key ingredients. They showed that the math boils down to two connected equations (or recipes) that are stuck together:
- The "Response Recipe": This tries to figure out who is missing the survey in the Control group. (Are the missing people mostly Compliers or Noncompliers?)
- The "Outcome Recipe": This tries to figure out what the missing people's scores would have been.
The problem is that you can't solve the "Outcome Recipe" until you solve the "Response Recipe," but they are tangled together.
The Solution: A Flexible Template
The paper proposes a unifying template (a flexible toolkit) to untangle these recipes. The authors say you need to make two separate choices to solve the puzzle:
Choice 1: The "Principal Identification" Assumption
This is your guess about how the program works on the "Slackers" (Noncompliers).
- Option A (Exclusion Restriction): "The program has zero effect on people who don't do the work." (If you don't volunteer, the program does nothing for you).
- Option B (Principal Ignorability): "If we look at people with similar backgrounds, the 'Slackers' and 'Hard Workers' would have had the same happiness scores in the Control group."
Choice 2: The "Missingness" Assumption
This is your guess about why the surveys are missing.
- Option A (Stable Response): "The rate at which people drop out is the same in the Treatment group as it is in the Control group." (e.g., If 20% of Hard Workers drop out in the Treatment group, we assume 20% drop out in the Control group too).
- Option B (Response Ignorability): "In the Control group, the 'Slackers' and 'Hard Workers' are equally likely to fill out the survey."
The Magic Trick:
The authors discovered that if you choose Option B for both choices (Principal Ignorability + Response Ignorability), the "Ghost" problem disappears entirely. The missing data becomes simple and random, and you don't need to make any extra guesses about the missing people!
A Warning About Old Rules
The paper also found a flaw in some old math rules (called "Stable Response"). Sometimes, these old rules can lead to impossible math answers, like saying "20% of people are missing" when the math actually implies "-5% are missing" (which is impossible).
To fix this, the authors suggest a "Near-Stable" version. It's like a safety guardrail: if the math tries to give an impossible answer, the rule gently bumps it back to the nearest possible number (like 0% or 100%) instead of breaking the whole calculation.
The Real-World Test
The authors tested their new template on data from the Baltimore Experience Corps (a real program where seniors volunteer in schools). They ran the numbers using different combinations of their "Choice 1" and "Choice 2."
What they found:
- The results changed depending on which assumptions they picked (which is expected).
- However, when they used the "Principal Ignorability" assumption, the results were very stable and didn't need extra guesses about the missing data.
- They confirmed that their new "Near-Stable" rules prevented the impossible math errors found in older methods.
Summary
This paper gives researchers a flexible map for navigating the messy world of missing data and non-compliance. Instead of being stuck with one rigid rule, researchers can now mix and match different assumptions to see how robust their results are. Most importantly, they found a special combination of assumptions that makes the missing data problem vanish, and they fixed a safety hazard in the old math tools.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.