← Latest papers
📊 statistics

Marginal and conditional summary measures: transportability and compatibility across studies

This paper clarifies the distinct interpretations and properties of marginal versus conditional summary measures, demonstrating that their non-coincidence—even for collapsible measures under effect modification—poses significant challenges for transportability and evidence synthesis, often leading to bias if incompatible measures are naively pooled without access to individual patient data.

Original authors: Antonio Remiro-Azócar, David M. Phillippo, Nicky J. Welton, Sofia Dias, A. E. Ades, Anna Heath, Gianluca Baio

Published 2026-05-07
📖 5 min read🧠 Deep dive

Original authors: Antonio Remiro-Azócar, David M. Phillippo, Nicky J. Welton, Sofia Dias, A. E. Ades, Anna Heath, Gianluca Baio

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to decide which of two different cars (Treatment A and Treatment B) is better. You can't test them both at the same time, so you look at two separate reports: one comparing Car A to a standard sedan (Control), and another comparing Car B to that same standard sedan. To figure out which is better, you have to combine these reports.

This paper is a warning label for that process. It argues that when we combine these reports, we often make a critical mistake: we are comparing apples to oranges without realizing it.

Here is the breakdown of the paper's main points using simple analogies.

1. The Two Ways to Measure "Success"

The paper explains that there are two main ways to measure how well a treatment works, and they don't always give the same number.

  • The "Group Average" (Marginal): Imagine you take a photo of the entire crowd of people who took the drug and ask, "On average, how much better did the group do?" This is a Marginal measure. It looks at the big picture.
  • The "Specific Profile" (Conditional): Now, imagine you zoom in on a specific type of person in that crowd—say, "people who are 30 years old and smoke." You ask, "How much better did this specific group do?" This is a Conditional measure.

The Problem: In the real world, these two numbers are rarely the same. Just because the "average person" improved by 10 points doesn't mean the "30-year-old smoker" improved by 10 points.

2. The "Magic Lens" Analogy

The paper introduces a concept called Collapsibility. Think of this as a special lens you look through to measure the treatment effect.

  • Directly Collapsible (The Clear Lens): Some lenses (like measuring simple "Risk Difference" or "Mean Difference") are clear. If you look at the specific groups and then average them up, you get the exact same number as if you looked at the whole crowd at once.
    • Analogy: If you measure the height of every student in a class and average them, you get the same result as measuring the whole class together.
  • Non-Collapsible (The Distorting Lens): Other lenses (like "Odds Ratios" or "Hazard Ratios," often used for survival or binary outcomes) distort the image. Even if you look at specific groups, when you average them up, the number changes.
    • Analogy: Imagine a funhouse mirror. If you look at a specific person, they look tall. If you look at the whole group, the mirror distorts the average height in a way that doesn't match the sum of the individuals.

The Paper's Big Surprise: Most people think this distortion only happens with the "funhouse mirrors" (non-collapsible measures). The paper argues that even with clear lenses, if the treatment works differently for different types of people (effect modification), the "Group Average" and the "Specific Profile" will still disagree.

3. The "Average Person" Trap

The paper highlights a confusing term often used in statistics: "At the Mean."

  • Imagine a study where people are measured by their "smoking status" (Yes/No).
  • The "Average" smoker might be 0.5 (half a person who smokes).
  • The paper warns that calculating the effect for this "0.5 person" is nonsense. It's like trying to calculate the fuel efficiency of a car that is half a sedan and half a truck. It doesn't exist.
  • The paper insists we must be careful not to confuse the effect for the "Average Person" (who doesn't exist) with the "Average Effect" across all real people.

4. The "Indirect Comparison" Disaster

This is where the paper gets practical. In healthcare, we often have to compare Treatment A vs. Treatment B, but we only have data for A vs. Control and B vs. Control. We have to do an "Indirect Treatment Comparison."

The paper says: If you mix the wrong types of measurements, your conclusion will be biased.

  • Scenario: Study 1 (A vs. Control) reports a "Group Average" result. Study 2 (B vs. Control) reports a "Specific Profile" result (adjusted for age, gender, etc.).
  • The Mistake: If you simply subtract Study 2's number from Study 1's number to compare A and B, you are mixing apples and oranges.
  • The Consequence: You might conclude Treatment A is better when it's actually worse, or vice versa, simply because the math doesn't line up.

5. The Solution: Full Access or Careful Math

The paper suggests two ways to fix this:

  1. The Gold Standard: Have access to the raw data for every single patient in every study (Individual Patient Data). This allows statisticians to recalculate everything to match perfectly.
  2. The Careful Approach: If you only have summary numbers (like "100 people took the drug, 20 got better"), you must use very specific mathematical methods (like ML-NMR) that know how to translate between "Group Averages" and "Specific Profiles" without breaking the math.

Summary

The paper is a call to stop being lazy with statistics. When comparing medical treatments:

  • Don't assume that an average effect is the same as a specific effect.
  • Don't assume that different studies are reporting the same thing just because they are measuring the same disease.
  • Check your lenses: Make sure you are comparing "Group Averages" to "Group Averages," and "Specific Profiles" to "Specific Profiles." If you mix them, your medical decisions could be wrong.

The authors aren't saying one type of measurement is "better" than the other; they are saying you must know exactly which one you are using and ensure you aren't accidentally mixing them up.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →