A critical evaluation of longitudinal proportional effect models
This paper critically evaluates nonlinear longitudinal proportional effect models, demonstrating that their assumption of a fixed treatment effect over time leads to bias, inflated Type I error rates, and safety concerns regarding the detection of treatment harm, even when the assumption holds.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Picture: A "Magic" Formula That Isn't Magic
Imagine you are running a race to see if a new running shoe (the Active Treatment) helps people run faster than their old shoes (the Control Group).
Scientists have been excited about a new mathematical tool called the "Longitudinal Proportional Effect Model." The pitch is that this tool is a "super-charger." It claims to be more powerful than standard methods and gives a direct answer to a very specific question: "By what percentage did the new shoes improve the runner's speed?"
However, this paper by Donohue, Insel, and Langford is a critical warning label. They argue that this "super-charger" is actually broken. It doesn't just give better answers; it gives biased answers that trick you into thinking a treatment works better than it actually does, while hiding the fact that it might be harmful.
The Core Problem: The "Zero Point" Trap
To understand the bias, imagine a thermometer.
- Standard Method (The Ruler): Measures the difference in degrees. If the room goes from 70° to 75°, the change is +5°. If it goes from 0° to 5°, the change is still +5°. It doesn't matter where you start.
- The "Proportional" Method (The Percentage): Measures the change relative to the starting point.
- If you start at 70° and go to 75°, that's a small percentage gain.
- If you start at 0° and go to 5°, the math breaks. You can't calculate a percentage of zero.
- If you start at -10° (below zero) and go to -5°, the math gets weird and unstable.
The paper argues that in medical trials (especially for diseases like Alzheimer's), the "starting point" (the control group's average score) often hovers near zero or even goes negative (representing a decline in health). When the math tries to calculate a "percentage improvement" from a number near zero, it becomes unstable and starts hallucinating results.
The "Favoritism" Bias: A One-Way Street
The authors discovered that this model has a built-in favoritism.
Imagine a judge who is secretly rooting for the new shoe brand.
- If the new shoes help: The model exaggerates how much they helped. It makes the improvement look huge.
- If the new shoes hurt: The model struggles to see the damage. It acts like a blind spot, making it very hard to detect if the treatment is actually making things worse.
The paper shows that this bias happens even when the rules of the model are followed perfectly. It's not a mistake in how the scientists used the tool; the tool itself is flawed.
The "Group Labeling" Trick
The paper points out a strange quirk: The model's answer changes depending on which group you call "Group A" and which you call "Group B."
- If you label the New Treatment as the main group, the model biases the result to make the treatment look good.
- If you swapped the labels and made the Control Group the main group, the model would bias the result to make the control look good.
This is like a scale that tips to the left if you put the object on the left side, and tips to the right if you put it on the right side. A good scale should give the same weight regardless of where you put the object. Because this model changes its mind based on labeling, its results are unreliable.
The Simulation Results: The "93% Error"
The authors ran computer simulations (virtual clinical trials) to test this. They found terrifying results:
- When the control group's average score was close to zero, the model started rejecting the idea that "nothing happened" 93% of the time, even when there was no treatment effect at all.
- In a standard test, you expect to be wrong about 5% of the time (this is called Type I error). This model was wrong nearly 20 times more often than it should be.
- In many cases, the model acted like a one-sided test: It would only shout "It works!" but rarely shout "It hurts!"
The Safety Concern: Hiding Harm
The most dangerous part of this bias is safety.
In diseases like Alzheimer's, there is a real fear that a drug could make memory worse (treatment harm).
- Because this model is biased toward seeing improvement, it is very bad at seeing decline.
- The authors warn that using this model could lead to approving a drug that actually harms patients, simply because the math was too busy trying to find a "percentage improvement" to notice the damage.
The Conclusion: Don't Use It
The authors conclude that despite the promise of higher statistical power (the ability to find a signal in the noise), this model is not worth the risk.
- The "power" it claims to have is actually just bias masquerading as success.
- Standard, linear models (the "rulers") are safer, more honest, and don't change their answers based on how you label the groups.
- They recommend that scientists avoid using these proportional effect models, especially for safety reasons.
In short: The paper says this new mathematical tool is like a broken speedometer that always reads "fast" when you are driving a new car, but fails to warn you when you are driving off a cliff. It's better to use an old, reliable speedometer than a fancy one that lies to you.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.