A Law of Iterated Expectation Primer for Causal Inference
This paper provides a primer clarifying the relationship between the law of iterated expectation and the g-formula for causal inference, presenting both non-iterative and iterative forms of the formula and illustrating their application through progressively complex numerical examples.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Picture: Why Do We Need This?
Imagine you want to know if a specific medicine (let's call it "Tamoxifen") actually prevents breast cancer from coming back. In a perfect world, you could give the medicine to one group of people and a sugar pill to another group, then compare the results. This is a randomized trial.
But in the real world, we often only have observational data. We can't force people to take medicine; we just watch what they choose to do. The problem is that people who choose the medicine might be different from those who don't (maybe they are sicker, or have different genetics). These differences are called confounders. If we don't account for them, we might blame the medicine for something that was actually caused by the patient's underlying health.
This paper is a "primer" (a beginner's guide) on how to fix this math problem. It explains a specific mathematical tool called the Law of Iterated Expectation and shows how it helps us turn messy, real-world data into a clear answer about cause and effect.
The Core Concept: The "Law of Iterated Expectation"
Think of this law as a way to calculate a weighted average.
Imagine you are a school principal trying to find the average test score for the whole school.
- The Simple Way: You just add up every single student's score and divide by the number of students. This is the "marginal expectation."
- The "Iterated" Way: You realize the school has different grades (1st grade, 2nd grade, etc.). You calculate the average score for 1st graders, then the average for 2nd graders, and so on. Then, you take those grade-level averages and combine them, but you weight them by how many students are in each grade.
The Law of Iterated Expectation simply says: You get the same final answer whether you average everyone at once, or if you average the groups first and then combine the groups.
In the paper, the authors explain that this mathematical identity is the engine behind the g-formula, a famous tool used to figure out causal effects.
Two Ways to Drive the Same Car: NICE and ICE
The paper introduces two different ways to use this math to solve the causal problem. They are mathematically identical (they give the exact same answer), but they look at the data differently. The authors call them NICE and ICE.
1. NICE (Non-Iterative Conditional Expectation)
The Analogy: The "Recipe Book" approach.
Imagine you want to know the average height of all people in a city, but you only have data on people who wear red hats and blue hats.
- How NICE works: You look at the "Red Hat" group and calculate their average height. You look at the "Blue Hat" group and calculate their average height. Then, you look at the city census to see what percentage of people wear red hats vs. blue hats. Finally, you mix those two averages together using the census percentages as weights.
- In the paper: The authors show this using a simple example of Tamoxifen and lymph nodes. They calculate the recurrence rate for different groups and then "plug in" the numbers to get a final weighted average.
2. ICE (Iterative Conditional Expectation)
The Analogy: The "Prediction Machine" approach.
Imagine you are a weather forecaster. Instead of just averaging past data, you build a model that predicts the weather for every single day based on the conditions of that day.
- How ICE works: You take your data, run it through a model, and generate a "predicted outcome" for every single person in your dataset (as if they had all taken the medicine). Then, you just take the average of all those predictions.
- In the paper: The authors show that you can do this by creating a list of "what-if" predictions for every person and then averaging them up.
The Key Takeaway: Whether you do the "Recipe Book" (NICE) or the "Prediction Machine" (ICE), you end up with the same number. The paper proves that these two methods are just two different ways of writing the same mathematical sentence.
Getting More Complex: Time and Moving Parts
The paper doesn't stop at simple examples. It shows how this works when things get complicated:
More Variables: What if you have age, income, race, and sex all mixed together? The "Recipe Book" (NICE) becomes very hard to write out because there are too many combinations. The "Prediction Machine" (ICE) is much easier because you just let the computer handle the math.
Time-Varying Confounders: This is the hardest part. Imagine a scenario where:
- You take a drug at Time 1.
- That drug changes your health (a confounder) at Time 2.
- That new health status affects whether you take a second dose of the drug at Time 2.
- Finally, you look at the outcome at Time 3.
In this scenario, standard statistics fail because the "confounder" (your health) was changed by the treatment itself. The paper shows that the g-formula (using the Law of Iterated Expectation) is the only way to untangle this knot. It does this by working backwards:
- First, predict the outcome at the very end.
- Then, work backward to predict what happened at Time 2.
- Then, work backward to Time 1.
- Finally, average it all out.
The paper calls this "backward recursion." It's like solving a maze by starting at the exit and walking backward to the entrance.
What the Authors Actually Claim (and What They Don't)
- They DO claim: The Law of Iterated Expectation is the mathematical foundation that allows us to turn "what we observed" into "what would have happened" (causal effects).
- They DO claim: The NICE and ICE methods are mathematically equivalent. They represent the same thing, just written differently.
- They DO claim: In simple situations (time-fixed), both methods are easy. In complex situations (time-varying), the ICE method (working backward) is often easier to code and more robust against certain types of errors.
- They DO NOT claim: This paper provides new medical results, new clinical guidelines, or specific advice for doctors on how to treat patients. It is purely a guide on the mathematics and logic of how to analyze data.
- They DO NOT claim: One method is "better" than the other in a general sense; they are just different tools for the same job. However, they note that if your mathematical models are wrong, both methods can give wrong answers.
The Bottom Line
This paper is a translator. It takes a very dense, scary mathematical concept (the Law of Iterated Expectation) and explains how it acts as the bridge between raw data and causal truth. It shows researchers that whether they use a "weighted average" approach or a "step-by-step prediction" approach, they are using the same fundamental logic to answer the question: "What would have happened if we had done something different?"
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.