← Latest papers
📊 statistics

Bounding causal effects with an unknown mixture of informative and non-informative missingness

This paper proposes a framework for bounding causal effects under an unknown mixture of informative and non-informative missingness, developing influence-function-based estimators that achieve root-n convergence and asymptotic normality while accommodating user-specified sensitivity parameters and competing outcomes like death.

Original authors: Max Rubinstein, Denis Agniel, Larry Han, Marcela Horvitz-Lennon, Sharon-Lise Normand

Published 2026-03-19
📖 5 min read🧠 Deep dive

Original authors: Max Rubinstein, Denis Agniel, Larry Han, Marcela Horvitz-Lennon, Sharon-Lise Normand

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to solve a mystery: Does Drug A cure a disease better than Drug B?

You have data from thousands of patients. But there's a problem: many patients disappeared from the study before you could see if they got better or worse. Some moved away (random), but others might have stopped taking the drug because it made them feel terrible, or because their condition got so bad they went to a different hospital (not random).

This is the core problem of missing data.

The Problem: The "Silent" Dropouts

In statistics, we usually assume that missing data is like a coin flip—random and unrelated to the outcome. This is called "Missing At Random" (MAR).

  • The Reality: Often, missing data is not random. It's "Informative."
  • The Analogy: Imagine a race. If the slowest runners drop out because they are tired, the remaining runners look faster than they really are. If you only look at the finishers, you think the race was easy. But if the fastest runners dropped out because they got injured, the finishers look slower than they really are.

The authors of this paper realize that in the real world, we rarely know why people dropped out. It's usually a mixture: some left for random reasons (lost their phone), and some left because of the outcome (got too sick).

The Solution: Drawing a "Safety Net"

Instead of guessing exactly why people left (which is impossible), the authors propose a new way to think about it: Bounding.

Think of the true answer (the real effect of the drug) as a bird flying in a foggy sky. We can't see the bird, but we can build a giant, invisible cage around the sky where the bird must be.

  • The "Worst-Case" Cage: If we assume everyone who dropped out was actually doing the worst possible thing, how bad could the drug look?
  • The "Best-Case" Cage: If we assume everyone who dropped out was doing the best possible thing, how good could the drug look?

The result isn't a single number (like "Drug A is 10% better"). Instead, it's a range (e.g., "Drug A is between 5% worse and 15% better"). This range is honest. It admits we don't know everything, but it tells us the limits of what is possible.

The "Sensitivity" Knob

The authors add a clever feature: a Sensitivity Knob.
Imagine you are adjusting the volume on a radio.

  • Turn the knob to "Low": You assume very few people dropped out for bad reasons. Your range is narrow, and you might say, "Drug A is definitely better!"
  • Turn the knob to "High": You assume many people dropped out because the drug was dangerous. Your range widens. Suddenly, you might say, "Well, Drug A could be better, but it could also be worse."

This allows researchers to ask: "How bad does the missing data have to be to change my conclusion?"

  • If you have to turn the knob to an impossible setting (e.g., "99% of dropouts were dying") to change your mind, then your conclusion is robust.
  • If a tiny turn of the knob changes everything, your conclusion is fragile.

The "Competing Event" (The Death of the Outcome)

The paper also handles a tricky scenario: Death.
If a patient dies, they can't get diabetes. If you only look at the people who survived to see if they got diabetes, you are ignoring the people the drug might have killed.

  • The Analogy: Imagine a video game. If a player dies, they can't finish the level. If you only count the players who finished the level, you might think the game is easy. But maybe the "hard mode" killed everyone.
  • The authors propose a new way to measure success that counts "Death" as a failure alongside "Getting Diabetes." This gives a more honest picture of the total risk.

Real-World Example: Antipsychotics and Diabetes

The authors tested their method on real data about antipsychotic drugs (used for schizophrenia) and diabetes risk.

  • The Naive View: If you just ignore the missing data, it looks like some drugs (like Haloperidol) are great at preventing diabetes.
  • The "Safety Net" View: When they applied their bounds, they found that the "great" result only holds if the missing data isn't too suspicious.
  • The Tipping Point: They calculated that for the "good" result to be true, the missing data would have to be extremely unbalanced (e.g., only the healthiest people dropped out). Since that seems unlikely, they concluded that while the drug might help, the evidence isn't as strong as the simple numbers suggested.

Why This Matters

Most studies give you a single number and a "confidence interval" that assumes everything is perfect. This paper says: "Life isn't perfect. Missing data is messy."

By using bounds and sensitivity knobs, researchers can:

  1. Stop pretending they know more than they do.
  2. Show exactly how much their conclusions depend on their assumptions.
  3. Give doctors and patients a clearer, more honest picture of risk, even when the data is incomplete.

In short: Instead of guessing the exact answer in the dark, this paper builds a flashlight that shows you the walls of the room, so you know exactly how big the space is, even if you can't see the center.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →