← Latest papers
📊 statistics

Transportability of model-based estimands in evidence synthesis

This paper argues that for non-collapsible effect measures, purely prognostic variables can induce marginal effect heterogeneity even when individual-level treatment effects are homogeneous, necessitating careful covariate adjustment in evidence synthesis to ensure transportability while highlighting the advantages of using directly collapsible measures.

Original authors: Antonio Remiro-Azócar

Published 2026-05-07
📖 6 min read🧠 Deep dive

Original authors: Antonio Remiro-Azócar

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a judge trying to decide which of two new medicines, Medicine A and Medicine B, works better for the general public. You can't test them head-to-head in a single giant trial because they were tested in different countries at different times. Instead, you have to do a "roundabout" comparison:

  • Trial 1 tested A against a placebo.
  • Trial 2 tested B against a placebo.
  • You compare the results of both trials against the placebo to see how A and B stack up against each other.

This is called an indirect comparison. The paper by Antonio Remiro-Azócar argues that the standard way of doing this math is often flawed, not because the trials were bad, but because of a hidden mathematical trap involving how we measure "success."

Here is the breakdown of the paper's core arguments using simple analogies.

1. The Two Ways to Measure "Success"

The paper focuses on two different ways to measure how well a drug works:

  • The "Average Person" View (Marginal): "If we give this drug to everyone in the city, how much better will the average person feel?"
  • The "Specific Profile" View (Conditional): "If we give this drug to a 50-year-old male who smokes, how much better will he feel compared to a non-smoker?"

The paper argues that for health policy (deciding what to reimburse for the whole population), we need the "Average Person" view. However, the math used in most drug trials often calculates the "Specific Profile" view.

2. The Trap: The "Non-Collapsible" Measures

The paper introduces a concept called collapsibility. Think of this like a Lego tower.

  • Collapsible (The Good Lego): If you have a tower made of blocks, and you take it apart, the total height is just the sum of the individual blocks. If you know the height of every block, you know the height of the tower.
    • Examples: Mean difference (e.g., "The drug lowers blood pressure by 5 points").
    • Result: If the "Specific Profile" view is the same for everyone, the "Average Person" view is automatically the same. You don't need to worry about the mix of people in the trial.
  • Non-Collapsible (The Tricky Puzzle): Imagine a puzzle where the picture changes depending on how the pieces are arranged. Even if every single piece (individual patient) reacts the same way to the drug, the final picture (the average result) looks different if the pieces are arranged differently.
    • Examples: Odds Ratios and Hazard Ratios (commonly used for binary outcomes like "cured vs. not cured" or time-to-event).
    • Result: The "Average Person" view changes even if the "Specific Profile" view stays the same, simply because the mix of people in the trial is different.

3. The "Purely Prognostic" Variables

The paper highlights a specific type of patient characteristic called a "purely prognostic" variable.

  • Analogy: Imagine a race. Some runners are naturally fast (prognostic). Some runners are slow. The drug helps everyone run 10% faster.
    • The "natural speed" doesn't change how much the drug helps (no interaction).
    • However, if Trial 1 has mostly fast runners and Trial 2 has mostly slow runners, the average improvement looks different in the two trials, even though the drug helped everyone equally.
  • The Paper's Claim: For "Non-Collapsible" measures (like Odds Ratios), these "natural speed" differences do change the final average result. For "Collapsible" measures, they do not.

4. The Simulation: What Happened in the Lab?

The author ran computer simulations to test this. He created two fake trials with different mixes of patients and compared three methods of analysis:

  1. The "Bucher Method" (The Standard): Just compares the averages directly. No fancy math adjustments.
  2. MAIC (Matching-Adjusted): Tries to re-weight the data so the patient mixes look the same.
  3. G-Computation (Modeling): Uses a complex mathematical model to predict what would happen.

The Findings:

  • Scenario A: No Interaction (The Drug works the same for everyone).
    • If the measure is Collapsible (Mean Difference): The standard method works fine.
    • If the measure is Non-Collapsible (Odds Ratio): The standard method fails. It gives a biased answer just because the patient mixes were different, even though the drug worked the same for everyone. You must adjust for the patient mix, even if they aren't "effect modifiers."
  • Scenario B: Interaction (The Drug works better for some people).
    • If the measure is Directly Collapsible (Mean Difference): You only need to adjust for the people who react differently.
    • If the measure is Not Directly Collapsible (Odds Ratio or Risk Ratio): It gets messy. The result depends on the entire relationship between all the patient variables (how age, weight, and gender correlate with each other).
    • The Catch: The standard data we have from published trials usually only gives us "averages" (mean age, mean weight). It rarely tells us how these variables correlate (e.g., do older people in Trial 2 also tend to be heavier?).
    • The Result: If you only adjust for the averages but ignore the correlations, your math is still biased. You need to know the "shape" of the whole group, not just the average.

5. The Big Takeaway for Decision Makers

The paper concludes that current guidelines for comparing drugs are too simple.

  • Current Rule: "Only adjust for variables that change how the drug works (effect modifiers)."
  • New Reality:
    1. If you use Odds Ratios (very common), you must adjust for all patient differences, even those that don't change how the drug works, because the math itself distorts the average.
    2. If you use Odds Ratios or Risk Ratios in a complex world where drugs work differently for different people, you need to know the full relationship between patient variables (correlations), not just their averages.
    3. Since we rarely have the full "relationship" data for the target population, our current methods might be giving us biased answers, leading to wrong decisions about which drugs to fund.

In short: The paper warns that using the wrong mathematical "ruler" (non-collapsible measures) combined with incomplete data (only knowing averages, not correlations) creates a "fog" that hides the true effectiveness of a drug when comparing it across different populations. To see clearly, we need better rulers (collapsible measures) or much better data (full joint distributions).

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →