← Latest papers
📄 medicine

Schema-enforced large language model framework for sarcoma tumor board simulation: a methodology development and reproducibility pilot study

This study developed and validated a schema-enforced large language model framework that achieves high reproducibility and eliminates structural hallucinations in simulating sarcoma tumor board decisions, establishing a robust methodology for future concordance studies against historical multidisciplinary team outcomes.

Original authors: Tekoshin Ammo, Moritz Englich, Fabio Dos Santos Adrego, Maximilian Jacobi, Alida Wilckens, Philipp Kruppa, Bastian Bonaventura, Stefan Niemuth, Selina M. Weiler, Victoria Wachenfeld-Teschner, Anja M.
Published 2026-07-31
📖 6 min read🧠 Deep dive

Original authors: Tekoshin Ammo, Moritz Englich, Fabio Dos Santos Adrego, Maximilian Jacobi, Alida Wilckens, Philipp Kruppa, Bastian Bonaventura, Stefan Niemuth, Selina M. Weiler, Victoria Wachenfeld-Teschner, Anja M. Boos, Gerrit Freund

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a super-smart, incredibly well-read robot how to act like a team of expert doctors. These doctors, called a "tumor board," meet up to discuss tricky cancer cases and decide on the best treatment plan. They are the ultimate decision-makers, combining knowledge from surgeons, radiologists, and oncologists to save lives. But here's the catch: the robot you are trying to train is a Large Language Model (LLM). Think of an LLM as a digital encyclopedia that has read almost everything ever written, but it has a mischievous habit. Sometimes, when asked a specific medical question, it might confidently make up a fake medical study, invent a statistic that doesn't exist, or give a different answer if you ask the exact same question twice. This is called "hallucinating," and it's like a student who guesses the answer on a test because they are too scared to say "I don't know."

For doctors, this is a big problem. If a robot gives a treatment plan that changes every time you ask, or if it invents a rule that doesn't exist, no one can trust it to help save a patient. The big question in this corner of science is: Can we build a "guardrail" for these robots? Can we force them to stop guessing and make up facts, and instead give us the exact same, reliable answer every single time we ask? If we can do that, maybe these robots could become helpful assistants in the doctor's office, helping to double-check plans or organize complex information without the risk of making things up.


The Robot's New Rulebook

In this study, a team of researchers from the University Hospital Schleswig-Holstein in Germany decided to test if they could build such a guardrail for a robot helping with a very specific, very tricky type of cancer: sarcoma. Sarcomas are rare tumors that can grow in bones or soft tissues, and because there are so many different types, they are like a puzzle with pieces that don't always fit together easily. The researchers wanted to see if they could make a robot act like a tumor board for these cases without the robot getting confused or making things up.

To do this, they didn't just ask the robot, "What should we do?" in a normal chat. Instead, they built a strict "rulebook" for the robot to follow. Imagine you are playing a video game where you can't just type anything you want; you have to pick your answers from a specific menu. If you try to type something that isn't on the menu, the game simply won't let you press "enter." The researchers did something similar. They forced the robot to output its answers in a very specific, structured format called JSON (which is just a fancy way of organizing data like a digital form). They gave the robot a list of 21 specific questions it had to answer, like "What kind of surgery?" or "How much radiation?" and forced it to pick its answers from a pre-approved list of options. If the robot tried to make up a new option or invent a fake medical guideline, the system would reject it immediately.

The Big Test

The team took 51 real, but completely anonymous, sarcoma cases from their hospital records. These cases covered all sorts of different patients, ages, and tumor types. They asked the robot to solve the puzzle for each case, but they didn't just ask once. They asked the exact same question three times for every single case. This was to see if the robot would give the same answer every time, or if it would flip-flop like a coin toss.

The results were surprisingly good. When the robot was forced to use this strict rulebook, it gave the exact same answer for 93.9% of the decisions across all three tries. That means out of more than 1,000 specific decisions the robot had to make, it only changed its mind about 6% of the time. Even better, the robot stopped making up fake facts entirely. In the past, robots might have invented a fake medical study to support a claim, but in this test, not a single fake guideline or made-up statistic appeared. The "hallucinations" were gone.

Where the Robot Still Hesitates

However, the robot wasn't perfect. The few times it did change its mind weren't because it was confused or making things up. Instead, it happened in situations where even human doctors might disagree. For example, when deciding on the best way to rebuild a body part after surgery (reconstruction), the robot sometimes wavered between a "local flap" and a "pedicled flap." The researchers noted that this isn't a mistake; it's just that in real life, both options are often valid, and different surgeons might choose differently based on their own style. Similarly, the robot sometimes varied on the exact dose of radiation, but all the doses it picked were within the safe and accepted range for doctors.

One specific case, involving an elderly patient with a very slow-growing tumor, caused the most trouble. The robot couldn't decide between watching the tumor and operating on it. The researchers realized this wasn't a bug; it was a feature. In this specific situation, even a human team of doctors might argue about the best path forward. The robot was correctly showing that there is no single "right" answer here.

What This Means

The main takeaway from this study is that by putting these strict "guardrails" on the robot, the researchers made it much more reliable. They proved that if you force a robot to follow a structured format, it stops making up facts and starts giving consistent answers. This doesn't mean the robot is ready to replace doctors tomorrow. The study only tested if the robot could be consistent, not if its medical advice was actually the best advice. But it's a huge step forward. It shows that we can build a system where the robot is a stable, predictable tool that doesn't lie or flip-flop, which is the first step toward letting it help doctors make real decisions in the future. The researchers are now planning to test if these consistent answers actually match what human doctors would have decided, but for now, they have successfully taught the robot to play by the rules.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →