← Latest papers
📊 statistics

Treatment effect estimation by comparing observed and predicted outcomes: conditions for valid inference and practical illustration

This paper utilizes the potential outcomes framework and a practical case study to clarify the necessary conditions for valid average treatment effect estimation when comparing observed outcomes under a new treatment with predicted outcomes from pre-existing models, a method commonly used in model-based clinical evaluation in radiotherapy.

Original authors: Lotta M. Meijerink, Artuur M. Leeuwenberg, Jungyeon Choi, Bas B. L. Penning de Vries, Johannes A. Langendijk, Judith G. M. van Loon, Remi A. Nout, Karel G. M. Moons, Ewoud Schuit

Published 2026-07-22
📖 5 min read🧠 Deep dive

Original authors: Lotta M. Meijerink, Artuur M. Leeuwenberg, Jungyeon Choi, Bas B. L. Penning de Vries, Johannes A. Langendijk, Judith G. M. van Loon, Remi A. Nout, Karel G. M. Moons, Ewoud Schuit

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to solve a mystery, but you have a very strange rule: you can only interview the suspects who are currently in the room, and you are forbidden from ever talking to the people who stayed home. This is the daily struggle of medical researchers when they want to know if a brand-new, shiny treatment actually works. Usually, the gold standard for solving this mystery is a "Randomized Controlled Trial," where doctors flip a coin to decide who gets the new medicine and who gets the old one. But sometimes, flipping a coin isn't possible, ethical, or fast enough. Maybe a new technology is too expensive to test on everyone, or the disease is too urgent to wait for a long study.

So, researchers try a clever trick called "counterfactual prediction." It's like building a time machine, but instead of sending a person back in time, they send a prediction back. They take a group of patients who received the new treatment and ask, "What would have happened to these specific people if they had received the old, standard treatment instead?" To answer this, they use a "model"—a mathematical recipe built from old data of people who only got the standard treatment. They plug the new patients' details into this recipe to generate a "ghost outcome," a prediction of what their health would look like under the old system. Then, they compare the real results of the new treatment against these ghost predictions. If the new treatment looks better, great! But here's the catch: this trick only works if the recipe is perfect, the time machine doesn't glitch, and the "ghosts" are truly identical to the real people. If any of these conditions are shaky, the whole mystery could be a false lead.

This paper, written by a team of researchers from the Netherlands, steps in to tidy up the detective's toolkit. They take this intuitive method—comparing real outcomes of a new treatment against predicted outcomes of an old one—and give it a rigorous mathematical makeover. They don't just say, "It works if you're careful"; they lay out exactly what "careful" means. The authors formalize the approach using a framework called "potential outcomes," which is just a fancy way of saying "what could have happened." They identify five specific conditions that must be true for the comparison to be fair and the results to be trustworthy.

Think of these five conditions as the five rules of a magic trick. If you break even one, the rabbit might not appear, or worse, you might pull out a duck and claim it's a rabbit. The first rule is Transportability: The "recipe" (the model) must work just as well in the new group of patients as it did in the old group. If the new patients are secretly different in a way the recipe doesn't know about (like a hidden ingredient), the prediction will be wrong. The second rule is Ignorability of Treatment Assignment: The reason a patient got the new treatment shouldn't be linked to how sick they would have been with the old treatment, once you account for what you know about them. If doctors only gave the new treatment to the "lucky" patients who were already destined to do well, the model will get confused. The third rule is Consistency: The treatment must mean the same thing every time. If "Proton Therapy" means something slightly different in the new hospital than it did in the old data, the comparison is apples to oranges. The fourth rule is Positivity: The new patients must look like people the model has seen before. You can't use a recipe designed for adults to predict the outcome for a toddler if the model has never seen a toddler. The fifth rule is Correct Model Specification: The mathematical recipe itself must be the right shape. If the relationship between a patient's age and their health is curved, but the recipe assumes it's a straight line, the prediction will be off.

The authors illustrate these rules with a real-world example involving head and neck cancer patients. They looked at a group of people who received a new, high-tech radiation treatment called proton therapy to see if it reduced swallowing difficulties (dysphagia) compared to the standard photon therapy. They used data from 750 patients treated with photon therapy between 2007 and 2017 to build their prediction model. Then, they applied this model to 93 patients treated with proton therapy in 2018 and 2019. The results suggested that proton therapy reduced the risk of swallowing problems by about 22% (a difference of -0.22) compared to what would have happened with photon therapy. However, the paper emphasizes that this number is only trustworthy if all five rules were followed.

The researchers also show how to check if the rules were followed. For instance, they looked at the patients in the new group who did receive the standard photon therapy and checked if the model predicted their outcomes correctly. If the model guessed right for them, it gave them more confidence that the "ghost predictions" for the proton group were also accurate. In their example, the model was very good at predicting the photon group (a difference of 0.0), which was a good sign.

Ultimately, the paper concludes that while this method is a powerful tool for making decisions when we can't run a perfect trial, it is not a magic wand. It requires a systematic check of all five conditions, supported by domain experts and real-world data. The authors warn that if these conditions are violated, the results could be biased, making a treatment look better or worse than it really is. They suggest that researchers should be transparent about these assumptions and use extra data to test them whenever possible. It's a reminder that in science, even the most elegant shortcuts require a solid foundation to stand on.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →