← Latest papers
📊 statistics

Bayesian fusion forests for heterogeneous treatment effects on survival from randomised and real-world data

This paper introduces the Bayesian fusion forest, a nonparametric framework that combines randomized controlled trial and real-world data to estimate heterogeneous treatment effects on survival outcomes by relaxing unconfoundedness assumptions and modeling confounding bias, thereby demonstrating improved efficiency and clearer clinical insights in a study of HIV antiretroviral therapy.

Original authors: Tijn Jacobs, Stéphanie L. van der Pas, Wessel N. van Wieringen

Published 2026-08-03
📖 6 min read🧠 Deep dive

Original authors: Tijn Jacobs, Stéphanie L. van der Pas, Wessel N. van Wieringen

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to solve a mystery: "Does this new medicine actually work, and does it work better for some people than others?" In the world of medicine, the gold standard for solving this is a Randomized Controlled Trial (RCT). Think of this as a perfectly controlled science fair where scientists randomly assign patients to get either the new medicine or a placebo. Because the assignment is random, any difference in health outcomes can be blamed on the medicine, not on the patients' lifestyles or genetics. However, these trials are often small, short, and expensive, like a quick flash of lightning that illuminates a tiny patch of ground.

On the other side of the street, we have Real-World Data (RWD). This is like a massive, messy, decades-long surveillance video of everyone in a city. It has thousands of patients and follows them for years, offering a much bigger picture. But there's a catch: in the real world, people aren't randomly assigned to take medicine. Sick people might take it because they feel terrible, or healthy people might take it because they are rich and have better insurance. This creates confounding, a sneaky bias where it looks like the medicine is causing the results, but it's actually the patients' background doing the heavy lifting.

The big question scientists face is: How do we combine the perfect logic of the small trial with the massive size of the messy real-world data without letting the messiness ruin the logic? If we just mash them together, the bias from the real world might trick us. If we ignore the real world, we miss out on valuable long-term clues. This is the puzzle tackled in a new paper by Tijn Jacobs and colleagues, who propose a clever new mathematical tool called the Bayesian Fusion Forest.

The Detective's New Toolkit: The Bayesian Fusion Forest

The authors of this paper have built a digital detective, a "forest" of decision trees, designed to merge these two very different data sources. Imagine you are trying to predict how long a patient will survive. The researchers realized that survival time is made up of several ingredients: a baseline prognosis (how sick the patient was to begin with), the treatment effect (how much the medicine helps), and a confounding function (the hidden bias in the real-world data).

Their secret sauce is a method that treats the real-world data with a "confounding function." Think of this as a noise-canceling headphone for data. The real-world data is full of static (bias), but the model uses the clean signal from the small trial to figure out exactly what that static sounds like. Once it knows the pattern of the noise, it can subtract it out, leaving behind the true signal of the medicine's effect. This allows them to use the huge real-world dataset to sharpen their estimates without letting the bias distort the answer.

The model is called a "forest" because it uses Bayesian Additive Regression Trees (BART). If you imagine a single decision tree as a flowchart asking "Is the patient older than 50? Yes/No," a forest is a massive team of thousands of these flowcharts working together. Each tree looks at a tiny piece of the puzzle, and when you combine their votes, you get a highly accurate, flexible prediction that can handle complex, non-linear relationships without needing the researchers to guess the formula in advance.

What They Found: Clarity from Chaos

The researchers tested their new forest in two ways: first, with computer simulations where they knew the "true" answer, and second, with real data from HIV patients.

In their simulations, they created fake scenarios where the real-world data was heavily biased and the two data sources were very different from each other. They found that their Bayesian Fusion Forest was incredibly robust. Even when the bias in the real-world data was strong, the forest managed to stay unbiased, whereas methods that just ignored the bias got the answer completely wrong. Furthermore, by borrowing strength from the larger real-world dataset, their method produced much tighter, more precise estimates than looking at the trial alone. It was like getting a high-definition photo from a blurry, low-resolution source by using the sharp focus of a smaller, perfect photo as a guide.

They also applied this to a real-life case: studying the effect of combination antiretroviral therapy for HIV. They combined a famous clinical trial (ACTG 175) with a massive long-term cohort study (MACS). The trial alone was too short to see the full picture of how long patients lived, and the real-world data alone was too messy to trust.

The results were striking. The trial alone suggested a benefit but was somewhat inconclusive, with wide margins of error. The real-world data alone was too noisy to be sure. But when the Bayesian Fusion Forest combined them, it revealed a clear, powerful benefit: the combination therapy stretched the time patients could live without the disease progressing by a factor of 1.65. In plain English, patients on the combination therapy lived about 65% longer than those on the older, single-drug treatment.

Perhaps most importantly, the model showed that this benefit wasn't just an average; it applied to nearly every single patient. While the trial alone left many patients in a "maybe" zone, the fusion method gave a high degree of certainty that the treatment helped almost everyone. It also discovered that the benefit was slightly higher for patients with higher CD8 cell counts and for white patients, showing how the treatment effect varies across different groups.

Why This Matters

The paper doesn't claim to have solved every problem in medicine, but it offers a powerful new way to listen to the data. It shows that we don't have to choose between the "perfect but small" trial and the "messy but big" real-world data. By using a smart, flexible model that can "cancel out" the noise of the real world, we can get answers that are both precise and trustworthy.

In the HIV example, the method turned a vague suggestion into a clear, actionable insight, identifying a treatment benefit for nearly every patient where the trial alone had been unsure. This suggests that for other diseases, combining these data sources could help doctors personalize treatments faster and more accurately, ensuring that the right medicine reaches the right person at the right time. The authors' work essentially builds a bridge between the controlled world of clinical trials and the chaotic reality of everyday life, allowing us to learn from both without getting lost in the noise.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →