Lord's 'paradox' explained: the 50-year warning on the use of 'change scores' in observational data
This paper resolves Lord's 50-year-old paradox by demonstrating that neither change-score comparisons nor baseline-adjusted follow-up comparisons can reliably estimate causal effects in observational data due to distinct biases, thereby highlighting the critical need for well-defined research questions and advanced causal inference methods like g-methods.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Mystery: Two Statisticians, Two Answers
Imagine a university wants to know if the food in the dining hall makes students gain weight, or if boys and girls gain weight differently. They weigh everyone in September (the start of the term) and again in June (the end of the term).
They hire two statisticians to look at the data.
- Statistician A looks at the difference in weight. They see that, on average, boys and girls weighed the same in September and weighed the same in June. Their conclusion: "Nothing changed. The diet and sex don't matter."
- Statistician B looks at the final weight but adjusts for how heavy the students were to begin with. They see that for students who started at the same weight, the boys ended up heavier than the girls. Their conclusion: "Boys gained significantly more weight than girls, even when they started at the same size."
This is Lord's Paradox. How can two smart people look at the exact same numbers and come to completely opposite conclusions?
The Paper's Solution: It's Not a Paradox, It's a Trap
The authors of this paper argue that this isn't a magical mystery. It's actually a warning sign. Both statisticians are using methods that are flawed in observational data (data where people aren't randomly assigned to groups, like boys vs. girls).
The paper explains that "change" is a tricky concept. To understand it, imagine a variable (like weight) changing for three different reasons:
- The "Habit" Factor (Endogenous Change): Things that happen automatically because of where you started. If you are already heavy, you might naturally stay heavy or gain a little just by existing.
- The "Noise" Factor (Random Change): Random fluctuations, like weighing yourself on a wobbly scale or having a bad night's sleep.
- The "Real" Factor (Exogenous Change): The actual effect of something new, like the diet or the exposure we are studying.
The Goal: We want to measure the "Real" factor. We want to know: Did the diet actually cause the weight change, independent of how heavy they were to start with?
Why Statistician A Failed (The "Change Score" Trap)
Statistician A used a Change Score. This is like calculating: Final Weight - Starting Weight.
The Analogy: Imagine you are trying to see how much a plant grew. You measure the plant at the start and the end, then subtract the start from the end.
- The Problem: If the "exposure" (like being a boy) caused the starting weight to be higher, this method breaks.
- The Paper's Explanation: When you subtract the starting weight, you accidentally subtract the effect of the exposure itself. If being a boy makes you start heavier, and you subtract that starting weight, you are essentially "canceling out" the very thing you are trying to measure.
- The Result: Statistician A got a result of "zero change" because the math canceled out the real effect. The paper says this method is dangerous in observational studies because it often hides the truth or even gives the wrong answer (like saying a medicine works when it actually doesn't, or vice versa).
Why Statistician B Was Also Flawed (The "Adjustment" Trap)
Statistician B used ANCOVA (Follow-up conditional on baseline). This is like saying, "Let's compare boys and girls, but only if they started at the exact same weight."
The Analogy: Imagine you are trying to see if a new fertilizer works. You look at two plants that are the same height. But, your ruler is slightly bent (measurement error).
- The Problem: In real life, our "starting weight" measurements aren't perfect. They have "noise" (like the wobbly scale).
- The Paper's Explanation: When you try to adjust for a starting measurement that is slightly wrong, the math gets messy. It tends to overestimate the effect. It looks like the boys gained more weight than they actually did because the "noise" in the starting measurement messed up the adjustment.
- The Result: Statistician B found a big difference, but the paper suggests this difference might be exaggerated because of measurement errors and other hidden factors (like physical activity) that weren't accounted for.
The Verdict: A 50-Year Warning
The paper concludes that neither method is perfect for observational data.
- Change Scores (Statistician A) tend to underestimate or hide the effect, especially if the exposure affects the starting point.
- Adjustment Methods (Statistician B) tend to overestimate the effect due to measurement errors and hidden confounders.
The "Perfect" Scenario:
The paper notes that if this were a Randomized Controlled Trial (like a drug trial where people are randomly assigned to groups), both methods would work fine and give the same answer. The paradox only happens in "observational" settings where groups (like boys vs. girls) already exist and have pre-existing differences.
The Takeaway for Everyone
The main lesson of this 50-year-old puzzle is: Be very careful when analyzing "change" in real-world data.
- Define your question clearly: What exactly are you trying to measure?
- Don't trust simple subtraction: Just taking the difference between "before" and "after" is often misleading if the groups started differently.
- Don't trust simple adjustments: Just "adjusting" for the starting point isn't a magic fix if your measurements are imperfect or if there are hidden factors (like exercise habits) influencing the results.
The paper doesn't offer a single "magic bullet" solution for all situations, but it warns scientists that the two most common ways of looking at change are often biased in opposite directions, leading to the confusing "paradox" we see in Lord's example.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.