← Latest papers
📄 medicine

OncoSynth: Synthetic data generation for treatment effect estimation in oncology

OncoSynth is a novel, causally-aware diffusion-based framework that generates high-fidelity synthetic oncology cohorts to overcome data access restrictions, significantly improving the accuracy of both population- and patient-level treatment effect estimation compared to existing methods.

Original authors: Stefan Feuerriegel, Octavia-Andreea Ciora, Julian Welzel, Dennis Frauen, Maresa Schröder, Marie Brockschmidt, Harry Amad, Thomas Callender, Mihaela van der Schaar

Published 2026-09-07
📖 4 min read☕ Coffee break read

Original authors: Stefan Feuerriegel, Octavia-Andreea Ciora, Julian Welzel, Dennis Frauen, Maresa Schröder, Marie Brockschmidt, Harry Amad, Thomas Callender, Mihaela van der Schaar

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the fight against cancer, doctors increasingly rely on vast collections of patient records to understand which treatments work best for whom. These records, often called real-world data, contain details about a patient's age, tumor size, and medical history, alongside the specific therapies they received and how long they survived. However, a significant barrier stands in the way of using this information freely: strict privacy laws and ethical rules prevent researchers from sharing the actual names and details of real patients. This creates a dilemma where valuable medical insights remain locked away because the data cannot be moved between hospitals or research centers. To solve this, scientists have developed a method called synthetic data generation. This process involves using computers to create entirely new, fake patient records that look and behave statistically like the real ones. These artificial records allow researchers to run experiments and test ideas without ever exposing a single real person's private information.

The challenge with existing methods of creating these fake records is that they often treat the data like a simple list of numbers to be copied, ignoring the cause-and-effect relationships that govern real medicine. In a real hospital, a doctor decides on a treatment based on a patient's specific condition, and that treatment then influences the patient's survival. If a computer model generates the treatment and the survival outcome at the same time, it might accidentally learn that the outcome caused the treatment, a logical impossibility that leads to misleading conclusions about how well a drug works. This flaw means that while older methods can create realistic-looking patient profiles, they often fail when used to answer the most critical question: does this treatment actually help?

A team of researchers has introduced a new system called OncoSynth, designed specifically to fix this problem in the field of oncology. Instead of generating all the data at once, OncoSynth follows the natural timeline of a patient's journey. First, it creates a synthetic patient with specific characteristics, such as age and tumor type. Next, it uses those characteristics to decide which treatment that synthetic patient would logically receive, just as a real doctor would. Finally, it determines the patient's survival outcome based on the treatment they received and their initial condition. By strictly following this chronological order, the system preserves the causal chain of events, ensuring that the fake data reflects how treatments truly affect survival.

The researchers tested this approach using two massive datasets from the United States: one containing records of 37,128 patients with lung cancer and another with 17,046 patients with breast cancer. They compared OncoSynth against other leading methods for creating synthetic data. The results showed that OncoSynth produced fake patient groups that were not only statistically similar to the real ones but also far superior at preserving the relationships between treatment and survival. When the researchers used the fake data to estimate how effective a treatment was for the general population, the error rate dropped by up to 66 percent compared to other methods. When they tried to predict which specific individuals would benefit most from a therapy, the error rate fell by up to 58 percent.

This improvement is crucial because it means that researchers can now use these synthetic datasets to generate reliable evidence about cancer treatments without needing access to the original, private records. The study demonstrated that the artificial data could accurately reproduce complex medical patterns, such as how certain tumor sizes influence the choice of chemotherapy or how different treatment sequences affect long-term survival. For instance, in the lung cancer group, the system correctly identified that radiation therapy was associated with better survival outcomes, and in the breast cancer group, it accurately reflected the differences in survival between patients who received chemotherapy before surgery versus those who received it after.

The researchers emphasize that while the system is highly effective, it is not a magic replacement for real-world trials. The quality of the fake data depends entirely on the quality of the real data it learns from; if the original records are incomplete or biased, the synthetic versions will reflect those same issues. However, the study confirms that this new approach offers a powerful tool for precision medicine. By allowing scientists to analyze treatment effects in a safe, privacy-preserving environment, OncoSynth helps unlock the potential of large medical databases, enabling the generation of clinical evidence that could lead to better, more personalized care for cancer patients everywhere.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →