Reframing Population-Adjusted Indirect Comparisons as a Transportability Problem: An Estimand-Based Perspective and Implications for Health Technology Assessment
This paper reframes population-adjusted indirect comparisons as a transportability problem, demonstrating that marginal treatment effects derived from such comparisons are often population-dependent due to non-collapsibility and effect modification, thereby necessitating explicit assumptions and careful alignment of estimands to ensure their validity for health technology assessment decision-making.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the world of medicine, deciding which new treatment works best for a specific group of patients often relies on comparing results from different clinical trials. These trials are the gold standard for proving that a medicine works, but they are rarely perfect mirrors of the real world. A trial might enroll only young, healthy volunteers, while the people who actually need the medicine are older and have other health conditions. When researchers want to compare two new drugs that have never been tested head-to-head, they must stitch together evidence from separate studies. This process, known as an indirect comparison, becomes complicated when the patient groups in those studies look different. To fix this, scientists use statistical tools to adjust the data, essentially reweighting the results so that the groups look more alike. This adjustment is meant to answer a crucial question for health agencies: if we give this drug to the people who actually need it, how much better will they do compared to the alternative?
A new paper by Conor Chandler and Jack Ishak challenges a long-held belief about how these adjustments work. For years, health technology agencies have operated under the assumption that if researchers correctly identify the factors that change how a drug works—such as age or disease severity—and adjust for them, the resulting comparison is valid for any population. The authors argue that this is not always true. They show that even when the adjustment is done perfectly, the way the final result is calculated can make it specific to the original study group and useless for the broader population. Their work reveals that the mathematical nature of the measurement itself—whether it is a simple difference in averages or a more complex ratio—determines whether the results can be safely moved from one group of people to another.
The researchers approached this problem by reframing the entire process as a question of "transportability." Imagine you have a map drawn for a specific city, and you want to use that map to navigate a different city that looks similar. If the streets are laid out in the exact same way, the map works. But if the layout is different, even slightly, the map might lead you astray. In the context of drug trials, the "map" is the statistical estimate of how well a drug works, and the "cities" are the different patient populations. The authors demonstrated that for many common ways of measuring drug effectiveness, the map is not universal. It is tied to the specific layout of the original city.
The paper focuses on two main ways of measuring treatment effects. The first is a conditional effect, which looks at how a drug works for a patient with a specific set of characteristics, like a 60-year-old with mild disease. The second is a marginal effect, which asks what the average outcome would be if everyone in a whole population received the drug. Health agencies usually care about the marginal effect because they need to know the impact on the entire population they serve. The authors found that while conditional effects can often be transported from one population to another if the right adjustments are made, marginal effects are much trickier.
Through a series of detailed simulations, the researchers tested five different types of measurements used in medical research. They created virtual populations with varying characteristics and tried to transport the results from one group to another. They found that for simple measurements, like the average difference in blood pressure or a simple risk difference, the results held up. If the adjustment was done correctly, the result was the same regardless of the population's makeup. However, for more complex measurements that are standard in the field, such as hazard ratios (used for survival data) and odds ratios (used for binary outcomes like success or failure), the results broke down. Even when the researchers assumed that the factors influencing the drug's effectiveness were identical across all groups, the marginal results still changed depending on the population.
This finding directly contradicts the current guidance used by major health assessment bodies. These agencies typically invoke the "shared effect modifier" assumption (SEMA) to justify transferring results from one study population to another. This idea suggests that if the same factors affect how different drugs work, then the comparison between those drugs should be the same for everyone. However, the authors show that this assumption is not sufficient on its own. While SEMA may allow for the transport of conditional effects, it fails to guarantee that the average, population-wide results will remain stable when moved to a new group. The problem lies in the mathematics of the measurement itself. Non-collapsible measures, which include the widely used odds and hazard ratios, do not behave like simple averages. They are influenced by the entire distribution of the population, not just the average characteristics.
The researchers illustrated this with specific examples. When they looked at a simple difference in means, the results were stable across all simulated populations. But when they switched to a log odds ratio, a common metric for binary outcomes, the results shifted as the population changed, even though the underlying drug performance and the adjustment factors remained constant. Similarly, for survival data, they examined the restricted mean survival time, a measure that calculates the average time a patient survives. Even though this measure is mathematically "collapsible" in a simple sense, it failed to transport correctly because the way the effect was modeled did not align with the way the final result was calculated. The mismatch between the scale used to model the data and the scale used to report the result created a hidden dependency on the population.
The implications of this work are significant for how medical evidence is used to make decisions about funding and access to new treatments. The authors argue that researchers and health agencies must stop assuming that a successful adjustment automatically makes a result transportable. Instead, they must carefully consider what kind of effect they are trying to measure and whether the mathematical properties of that measure allow it to be moved across populations. If the goal is to estimate a population-wide effect, using a non-collapsible measure like an odds ratio may lead to biased conclusions if the results are simply applied to a new group without further adjustment.
The paper does not suggest that indirect comparisons are useless. Rather, it calls for a more precise and honest approach to how these comparisons are interpreted. It suggests that when a study produces a result that is inherently tied to the specific population it was measured in, that result should be treated as population-specific. It should not be blindly applied to a different group of patients, even if the groups look similar on paper. For health technology assessments, this means that the choice of statistical method and the type of effect measure are not just technical details; they are fundamental to whether the evidence can be trusted for real-world decision-making.
Ultimately, the work provides a clear framework for determining when a treatment effect can be safely transported and when it cannot. It shows that direct transportability is not a given; it is a property that depends on the alignment of the statistical model, the type of effect being measured, and the assumptions about how the population is structured. By distinguishing between conditional and marginal effects and understanding the role of collapsibility, researchers can avoid the trap of assuming that a corrected result is universally applicable. The path forward involves recognizing that some results are local to the study that generated them, and that applying them elsewhere requires additional, often unspoken, assumptions that may not hold true. This clarity allows for more transparent and principled use of indirect evidence, ensuring that decisions about patient care are based on a solid understanding of what the data can and cannot tell us.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.