← Latest papers
🤖 machine learning

OncoSynth: Synthetic data generation for treatment effect estimation in oncology

OncoSynth is a novel, causally-aware diffusion-based framework that generates high-fidelity synthetic oncology cohorts to overcome data privacy restrictions, significantly improving the accuracy of both population- and patient-level treatment effect estimation compared to existing methods.

Original authors: Octavia-Andreea Ciora, Julian Welzel, Dennis Frauen, Maresa Schröder, Marie Brockschmidt, Harry Amad, Thomas Callender, Mihaela van der Schaar, Stefan Feuerriegel

Published 2026-06-25
📖 4 min read☕ Coffee break read

Original authors: Octavia-Andreea Ciora, Julian Welzel, Dennis Frauen, Maresa Schröder, Marie Brockschmidt, Harry Amad, Thomas Callender, Mihaela van der Schaar, Stefan Feuerriegel

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a chef trying to perfect a new recipe for a life-saving medicine. To do this, you need to test it on thousands of patients to see who gets better and who doesn't. However, there's a huge problem: you can't actually use real patients for these tests because their medical records are private, locked away for legal and ethical reasons. It's like trying to learn how to bake a cake without ever being allowed to touch the real ingredients or see the real bakers.

This is where OncoSynth comes in. Think of OncoSynth as a "digital twin" factory that builds a perfect, fake version of a patient population. But here's the catch: most existing factories make "fake" patients that look real on the surface (they have the right age, weight, and tumor size) but fail to understand the story of how they got sick and how they were treated. They mix up the timeline, like a movie where the ending is shown before the beginning, which leads to wrong conclusions about whether a treatment actually works.

The Problem with Old Methods

Imagine you are trying to learn how a specific type of rain (treatment) affects a garden (patient outcome).

  • Old methods are like taking a snapshot of the garden, the rain, and the flowers all at once and shoving them into a blender. The blender spits out a smoothie that looks like a garden, but it has lost the logic that rain causes flowers to grow. Because the blender mixed everything together, it might accidentally learn that "flowers cause rain," which is impossible. In medical terms, this leads to biased results where a drug might look effective when it's not, or vice versa.

How OncoSynth Works

OncoSynth is different because it acts like a chronological storyteller rather than a blender. It builds the fake patients step-by-step, strictly following the order of real life:

  1. Step 1: The Patient Profile. First, it generates a patient's background (age, genetics, tumor size). This is the "before" picture.
  2. Step 2: The Doctor's Decision. Next, it decides what treatment that specific patient would have received based on their profile. It asks, "Given this patient's age and tumor, what would a real doctor have prescribed?"
  3. Step 3: The Outcome. Finally, it simulates what happens to that patient after receiving that specific treatment. It asks, "Given this patient and this treatment, how long do they survive?"

By keeping these steps separate and in order, OncoSynth ensures that the "fake" patients have the same logical cause-and-effect relationships as real people. The treatment doesn't magically change the patient's age, and the outcome doesn't influence the doctor's past decision.

The Results: A Better "Fake" World

The researchers tested this system using real data from two massive groups of patients: one with lung cancer (over 37,000 people) and one with breast cancer (over 17,000 people). They compared OncoSynth against the best existing "blender" methods.

  • The Look: OncoSynth created fake patients that looked statistically identical to real ones. The distribution of ages, tumor sizes, and survival times matched the real world almost perfectly.
  • The Logic (The Big Win): The real magic happened when they tried to measure how well the fake data could predict treatment success.
    • When estimating the average benefit of a treatment for a whole population, OncoSynth reduced the error by up to 66% compared to other methods.
    • When estimating the personalized benefit for a specific individual (precision medicine), it reduced the error by up to 58%.

Why This Matters

In the world of oncology (cancer research), knowing exactly how a treatment works for different types of people is crucial. But because real data is so hard to share, researchers often can't run these tests.

OncoSynth provides a solution: a safe, synthetic playground. It allows researchers to generate a "fake" dataset that is so faithful to reality—not just in appearance, but in the causal logic of how treatments work—that they can run their experiments, figure out which drugs work best for whom, and generate reliable evidence, all without ever seeing a single real patient's private record.

In short, OncoSynth doesn't just make fake patients; it makes fake patients that think and react like real ones, allowing scientists to learn how to save lives without breaking privacy rules.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →