← Latest papers
📊 statistics

Self-separated and self-connected models for mediator and outcome missingness in mediation analysis

This paper addresses identification challenges in mediation analysis with missing mediators and outcomes by introducing and synthesizing self-separated and self-connected missingness models, the latter of which leverages shadow variables and unobserved dependencies to provide robust theoretical templates and practical extensions for causal inference.

Original authors: Trang Quynh Nguyen, Razieh Nabi, Fan Yang, Grace V. Ringlein, Elizabeth A. Stuart

Published 2026-04-07
📖 5 min read🧠 Deep dive

Original authors: Trang Quynh Nguyen, Razieh Nabi, Fan Yang, Grace V. Ringlein, Elizabeth A. Stuart

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to figure out why a specific intervention works. You have three main characters in your story:

  1. The Treatment (A): The thing you did (like a job training program or a school anti-drinking campaign).
  2. The Mediator (M): The middleman (like a student's attitude or a person's education level).
  3. The Outcome (Y): The final result (like whether they got a job or stopped drinking).

Usually, researchers want to know: Did the treatment work because it changed the middleman, which then changed the result?

The Problem: The Missing Pages
In the real world, data is messy. Sometimes people don't show up for the survey, or they refuse to answer specific questions.

  • The Mediator is missing: We don't know their attitude or education level.
  • The Outcome is missing: We don't know if they got a job or stopped drinking.

The biggest headache happens when both are missing, and the reason they are missing is related to the missing data itself. This is called "Not Missing At Random" (MNAR).

  • Example: A student who is drinking heavily might be too embarrassed to answer the question about drinking. Or, someone who failed to get a job might be too discouraged to fill out the survey.
  • The Trap: If you just ignore the missing people (or guess they are like the people who answered), your results will be wrong. You might think the program worked, when actually it only worked for the people who were willing to talk to you.

The Paper's Solution: Two Ways to Fix the Puzzle

This paper is like a detective's guidebook for solving these "missing page" mysteries. The authors propose two main strategies to figure out the truth without needing to magically fill in the blanks.

Strategy 1: The "Self-Separated" Models (The "Clean Break" Approach)

Imagine you are trying to guess a secret password.

  • The Rule: In this approach, we assume that the reason a page is missing is completely unrelated to the content of the page, once we account for everything else we know (like age, income, or school).
  • The Metaphor: Think of a library where books go missing. If the books go missing because the librarian is lazy (a random cause), and not because the books are boring (the content), we can still figure out the story of the library.
  • The Catch: This only works if we can prove that the missingness isn't caused by the variable itself. If the "boring books" are the ones getting lost, this strategy fails. The paper shows that while this is a common assumption, it's often too strict and unrealistic for sensitive topics (like drinking or unemployment).

Strategy 2: The "Self-Connected" Models (The "Shadow Variable" Approach)

This is the paper's superpower. It acknowledges that sometimes, the missingness is caused by the variable itself (e.g., heavy drinkers hide their drinking). But, we can still solve the puzzle if we have a Shadow Variable.

  • What is a Shadow Variable?
    Imagine you are trying to guess a person's height (the missing data), but they are standing behind a curtain. You can't see them. However, you have a Shadow Variable: a very accurate photo of their shadow cast on the wall.

    • The shadow (Shadow Variable) is linked to the person's height.
    • But the shadow is not affected by the person's decision to hide behind the curtain (the missingness).
    • By studying the shadow, you can mathematically reconstruct the person's height, even though you can't see them.
  • Where do these Shadows come from?
    The paper explains how to find these "shadows" in real life:

    1. Built-in Shadows: Sometimes the Mediator or Outcome itself can act as a shadow for the other. (e.g., If we know a student's attitude from a different time, it can help us guess their drinking behavior, even if the drinking data is missing).
    2. External Shadows: We can use other data we collected.
      • Example: In a job training study, if we are missing "earnings" data, maybe we have "home ownership" data or "credit scores" from a later date. These are linked to earnings but aren't the earnings themselves, so they act as a shadow.
      • Example: In a school study, if parents didn't report on their child's behavior, maybe the teachers' reports on the same child can act as a shadow.

The "Time Travel" Factor

The paper also gets very clever about time.

  • Did the missing data happen before or after the event?
  • If you ask about past behavior (retrospective), the rules change.
  • The authors map out different "time travel" scenarios to show which Shadow Variable strategy works best depending on when the data was collected.

Why This Matters

Most standard statistics software just assumes missing data is random (like a coin flip). This paper says, "Stop assuming that!"

It provides a toolkit for researchers to:

  1. Check if their model makes sense: It gives them "testable clues" (like checking if the shadow variable behaves the way it should) to see if their assumptions are wrong.
  2. Use extra data: It encourages researchers to collect "shadow" data (like teacher ratings, bank records, or past surveys) specifically to help solve the missing data problem.
  3. Get the truth: By using these advanced models, we can stop guessing and start calculating the true effect of interventions, even when people hide the truth or drop out of studies.

In a Nutshell:
If your data is missing because the answer is "embarrassing" or "bad," you can't just ignore it. You need a Shadow Variable—a clue that is related to the answer but not affected by the embarrassment. This paper teaches you how to find those clues and use them to solve the mystery of the missing data.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →