← Latest papers
📊 statistics

Hidden in Plain Sight: How Non-Collapsibility Biases Treatment Effects in (Network) Meta-Analysis

This paper demonstrates that standard network meta-analysis models produce biased, null-ward estimates for non-collapsible effect measures like the odds ratio when studies involve heterogeneous populations with varying baseline risks, and proposes a "bookend" approach to explicitly model these mixed populations as weighted combinations of homogeneous subgroups to correct the bias.

Original authors: Harlan Campbell, Jeroen P. Jansen

Published 2026-03-03
📖 5 min read🧠 Deep dive

Original authors: Harlan Campbell, Jeroen P. Jansen

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a chef trying to figure out which of two new spices (let's call them Spice A and Spice B) makes a soup taste better. You don't have one giant pot of soup to test them on; instead, you have to look at recipes from three different kitchens.

Here is the problem: The "Standard Recipe" for comparing these spices is secretly broken.

The Setup: Three Kitchens, One Goal

Let's say you are looking at data from three different studies (kitchens) to see if Spice B is better than Spice A for preventing a specific flavor defect (the "event").

  1. Kitchen 1 (The High-Risk Group): This kitchen only uses ingredients that are naturally very salty. Their soup is always salty, no matter what.
  2. Kitchen 2 (The Low-Risk Group): This kitchen only uses ingredients that are naturally very bland. Their soup is always bland, no matter what.
  3. Kitchen 3 (The Mixed Group): This kitchen is a chaotic mix. Half the time they use salty ingredients, half the time they use bland ones.

The Truth: In reality, Spice B is equally amazing in both kitchens. It reduces the "saltiness defect" by exactly 50% in the salty kitchen and by exactly 50% in the bland kitchen. The spice works perfectly the same way for everyone.

The Trap: The "Average" Illusion

Now, imagine you are the statistician trying to combine these results. You use the Standard Model (the usual way scientists do this).

  • Kitchen 1 reports: "Spice B cut the defect by 50%." (Great!)
  • Kitchen 2 reports: "Spice B cut the defect by 50%." (Great!)
  • Kitchen 3 reports: "Spice B cut the defect by... only 40%?"

Wait, why did Kitchen 3 report a weaker result if the spice works the same?

This is the magic trick of Non-Collapsibility.

Think of it like a volume knob.

  • In the Salty Kitchen, the volume is already turned up to 10. Turning it down by 50% makes a huge, noticeable difference.
  • In the Bland Kitchen, the volume is at 1. Turning it down by 50% is barely noticeable.
  • In the Mixed Kitchen, you have a crowd of people. Some have the volume at 10, some at 1. When you turn the "Spice B" knob down, the people at volume 10 drop to 5 (a big change), but the people at volume 1 drop to 0.5 (a tiny change).

When you calculate the average result for the whole mixed crowd, the math gets messy. Because the "baseline" (how salty the soup started) was so different, the average improvement looks smaller than it actually is. The result gets pulled toward the middle (the "null"), making the spice look less effective than it really is.

This is what the paper calls bias toward the null. The standard model sees the Mixed Kitchen's "40%" result and averages it with the others, concluding: "Spice B is good, but maybe only 45% effective."

The tragedy: The spice is actually 50% effective, but the math tricked you into thinking it was weaker.

The Solution: The "Bookend" Approach

The authors, Harlan and Jeroen, propose a clever new way to fix this, which they call the "Bookend" Model.

Imagine a bookshelf.

  • On the far left, you have a book representing the Salty Kitchen (Extreme High Risk).
  • On the far right, you have a book representing the Bland Kitchen (Extreme Low Risk).
  • In the middle, you have the Mixed Kitchen.

The Standard Model tries to guess the middle by just averaging the numbers.
The Bookend Model says: "Wait a minute. The Mixed Kitchen is just a copy of the Left Book and the Right Book glued together."

Instead of treating the Mixed Kitchen as a mystery, the model assumes:

  1. The "Salty" study is 100% Salty people.
  2. The "Bland" study is 100% Bland people.
  3. The "Mixed" study is just a 50/50 mix of those two groups.

By mathematically "un-gluing" the Mixed Kitchen back into its two pure parts, the model can see that the spice is actually doing its full 50% job in both parts. It corrects the bias and tells you the truth: Spice B is 50% effective.

Why Should You Care?

This isn't just about soup spices. This happens in medicine all the time.

  • Scenario: Doctors want to know if a new cancer drug works.
  • The Trap: Some studies test the drug on very sick patients (high risk of death). Others test it on healthy patients (low risk).
  • The Result: If you use the standard math, the drug might look less effective than it really is because the "mixed" studies dilute the results.
  • The Danger: A doctor might decide not to prescribe a life-saving drug because the math made it look "meh," when in reality, it's a miracle cure for specific groups.

The Takeaway

The paper warns us that averaging averages can be dangerous when the groups being averaged are fundamentally different.

  • The Old Way: "Let's just take the average of all the studies." (This hides the truth).
  • The New Way (Bookend): "Let's find the most extreme studies (the bookends), assume they are pure groups, and figure out how the mixed studies are built from them."

It's a reminder that in science, context matters. You can't just mix apples and oranges and expect the average to tell you what an apple tastes like. You have to understand the ingredients before you can judge the recipe.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →