← Latest papers
📊 statistics

Effect modification and non-collapsibility leads to conflicting treatment decisions: a review of marginal and conditional estimands and recommendations for decision-making

This paper reviews how the combined presence of effect modification and non-collapsibility can cause marginal and conditional estimands to yield conflicting treatment rankings, and it provides practical recommendations for decision-making that leverage multilevel network meta-regression to generate both types of estimates for the target population.

Original authors: David M. Phillippo, Antonio Remiro-Azócar, Anna Heath, Gianluca Baio, Sofia Dias, A. E. Ades, Nicky J. Welton

Published 2026-08-18
📖 5 min read🧠 Deep dive

Original authors: David M. Phillippo, Antonio Remiro-Azócar, Anna Heath, Gianluca Baio, Sofia Dias, A. E. Ades, Nicky J. Welton

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the world of healthcare, doctors and policymakers face a constant challenge: deciding which treatment works best for a specific group of people. They rely on clinical trials, which are carefully controlled experiments, to find the answers. However, people are not all the same. Some have severe forms of a disease, others have mild forms, and some carry specific biological markers that change how their bodies react to medicine. When a factor like this changes how well a treatment works, scientists call it "effect modification." It is a well-known fact that if a drug works wonders for one type of patient but barely helps another, the best choice for the whole group depends on who is in that group.

To make these decisions, researchers often combine data from multiple trials to get a clearer picture. But there is a mathematical quirk that complicates things. When measuring the success of a treatment, scientists often use ratios, such as odds ratios or hazard ratios, which compare the likelihood of an event happening in one group versus another. These measures are "non-collapsible," a technical way of saying that the average result for a whole group is not simply a weighted average of the results for the individuals inside it. Because of this, the way you calculate the average matters. You can calculate the average by looking at the effect on each person and then averaging those effects, or you can calculate the average by looking at the overall event rates for the whole group and comparing them. Usually, these two methods give slightly different numbers, but they generally agree on which treatment is better.

A new paper by David Phillippo and his colleagues reveals that this agreement is not guaranteed. The researchers show that when effect modification is present, these two ways of calculating the average can lead to completely opposite conclusions. One method might suggest that Treatment A is the best choice for the entire population, while the other method suggests that Treatment B is superior. This happens because the two methods are actually answering two different questions. One asks, "Which treatment gives the best average benefit to the individual patients?" while the other asks, "Which treatment results in the fewest total bad events for the entire population?" When the effectiveness of a treatment varies significantly across different types of patients, the answer to these two questions can diverge.

The authors illustrate this conflict using a series of simulations. In one scenario, they modeled a population where a specific biomarker made a treatment highly effective for some people but less effective for others. When they calculated the average benefit for the individual, the data pointed to one drug as the clear winner. However, when they calculated the overall event rate for the whole group, a different drug appeared to be the best choice. In this specific example, choosing the drug that minimized the total number of bad events meant that 75 percent of the population would receive a treatment that was actually inferior for them personally. Conversely, choosing the drug that was best for the average individual would result in a higher total number of bad events across the population.

This finding challenges the common assumption that there is always a single "best" treatment for a population. The researchers argue that the conflict arises because the mathematical curves representing the probability of an event for different treatments can cross each other. If a treatment is better for one subgroup but worse for another, and these subgroups are mixed together, the overall average can flip the ranking. This is particularly relevant for time-to-event outcomes, such as survival times, where the presence of any patient characteristic, even one that does not change the treatment effect, can cause the average risk to change over time.

The paper also examines the tools currently used to adjust for these differences between study populations. Methods like matching-adjusted indirect comparison and simulated treatment comparison are widely used to tailor trial results to a specific target population. However, the authors demonstrate that these methods often produce estimates that are specific only to the population of the original study and cannot be easily transported to a new population with different characteristics. They identify a more advanced method, multilevel network meta-regression, as the only current approach capable of producing both types of estimates—those for the individual and those for the group—in any target population.

Ultimately, the authors recommend that decision-makers stop relying on a single number to make choices. Instead, they should look at both the individual-level and the population-level estimates. If both methods agree on the best treatment, the decision is straightforward. If they disagree, it signals that the population is not uniform and that a single decision for everyone may be suboptimal. In such cases, the best course of action might be to split the population into subgroups and assign different treatments to each group. This approach ensures that patients receive the treatment that is best for their specific characteristics, rather than forcing a compromise that might leave the majority worse off. The paper concludes that while identifying these subgroups is difficult and requires careful statistical work, it is the only way to resolve the conflict between maximizing individual benefit and minimizing total harm when effect modification is present.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →