Treatment effect: a critique
This paper critiques the distinction between model-based and counterfactual definitions of treatment effects, arguing that while counterfactual frameworks offer idealized clarity, model-free definitions may be unsuitable for generalizing scientific conclusions beyond the observed sample, echoing concerns raised by Cox.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a chef trying to figure out if adding a secret spice (the "treatment") makes a soup taste better. You want to know exactly how much better it is. This paper is a debate between two groups of statisticians about the best way to measure that "betterness."
Here is the breakdown of the paper's argument, using simple analogies.
The Two Competing Recipes
The authors, Battey and Edgar, are comparing two different ways to define a "treatment effect" (the impact of the spice).
1. The "Model-Based" Recipe (The Fisherian Approach)
- The Idea: This approach assumes there is a hidden, stable rule governing how the soup changes. Think of it like a physics law. If you add the spice, the soup's flavor shifts by a specific, consistent amount (like adding exactly 2 grams of salt).
- The Goal: To find that stable "rule" or parameter. Even if you cook the soup with different base ingredients (different people), the change caused by the spice remains constant according to the rule.
- The Analogy: Imagine a conveyor belt where every box gets a sticker added. The "treatment effect" is the size of the sticker. It's the same size for every box, regardless of what's inside the box.
2. The "Model-Free" Recipe (The Counterfactual Approach)
- The Idea: This approach doesn't assume a hidden rule. Instead, it asks a "What if?" question for every single person. "What would Person A have tasted if they didn't get the spice, compared to what they did taste?"
- The Goal: To calculate the average difference between the "real world" and the "imaginary world" for everyone in the group.
- The Analogy: Imagine you have 100 clones of yourself. You give the spice to one clone and not the other. Then you do it again for a different set of clones. You take the average of all those differences. This is the "Model-Free" method.
The Big Problem: The "Chameleon" Effect
The authors argue that the "Model-Free" approach has a major flaw: It changes depending on who you ask.
To prove this, they use a thought experiment (a "fictitious idealization"). They pretend we have a magic machine that lets us see both the "spiced" and "unspiced" versions of every person at the same time. Even with this perfect information, the Model-Free average behaves strangely.
The Analogy of the Unstable Ruler:
Imagine you are measuring the height of a group of people to see how much a growth hormone (the treatment) helps them.
- The Model-Based view says: "The hormone makes everyone grow exactly 2 inches." This is a stable fact.
- The Model-Free view calculates the average growth.
- If your group consists of very short people, the hormone might make them grow 2 inches, which is a huge percentage increase.
- If your group consists of very tall people, the same 2 inches might be a tiny percentage increase.
- The Result: The "average effect" calculated by the Model-Free method changes completely just because you picked a different group of people to measure.
The authors show that in many real-world scenarios (like binary outcomes like "sick" vs. "healthy"), the Model-Free average depends entirely on the specific mix of people in your sample. If you change the sample, the "answer" changes.
Why This Matters
The authors argue that if you want to make scientific conclusions that apply to the world at large (not just the specific people in your study), you need a stable definition.
- The Critique: The Model-Free approach is like trying to measure the "average speed" of a car by looking at how fast it goes on different roads. If the road is steep, it goes slow; if it's flat, it goes fast. The "speed" isn't a property of the car; it's a property of the specific road you happened to drive on.
- The Consequence: If you use the Model-Free method, your conclusion about the treatment might be true for your specific group of patients, but it might be completely wrong for the next group of patients you see. It doesn't generalize well.
Addressing the Counter-Arguments
The paper also tackles three common defenses of the Model-Free approach:
"But people are different!"
- Defense: The Model-Free method is better because it lets different people have different effects.
- Rebuttal: The Model-Based method can handle this too! It just treats the differences as "interactions" (like saying the spice works differently on spicy soup vs. bland soup) while keeping the core rule stable.
"We don't want to assume a model!"
- Defense: We shouldn't guess the rule; we should just measure the difference.
- Rebuttal: If you don't assume a rule, you lose the ability to say anything stable about the future. You end up with a number that only works for the specific people you measured today.
"It has 'Double Robustness'!"
- Defense: The math is clever; it works even if one part of the calculation is wrong.
- Rebuttal: This is only useful if the thing you are measuring (the average difference) is actually a meaningful, stable target. If the target itself is wobbly (like the unstable ruler), being "robust" doesn't help you hit the bullseye.
The Bottom Line
The authors conclude that while the "Model-Free" (counterfactual) approach is popular and mathematically clever, it often fails to provide a stable definition of a treatment effect that can be generalized to new people.
They argue that we should return to the "Model-Based" approach (associated with R.A. Fisher), where we look for a stable rule or parameter that explains how the treatment works, rather than just averaging up the differences in a specific, potentially unrepresentative group.
In short: Don't just average the differences you see in your specific group. Find the underlying rule that explains the change, so your conclusions hold true for the next group you meet.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.