Agentic systems for breast cancer treatment recommendations
This study evaluates agentic large language model systems for breast cancer treatment recommendations using 72 real clinical cases and 1,147 asymmetric rubrics, finding that while the best-performing configuration achieved a global score of 0.594, persistent clinically relevant failures indicate these systems are currently insufficient for unsupervised clinical use.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you're trying to build the ultimate "Medical Super-Brain" to help doctors decide the best way to treat breast cancer. You have a library of 72 real-life patient stories, ranging from early-stage cases to advanced ones, and you want to see if a team of AI robots can read these stories and write a perfect treatment plan. That's exactly what this study did.
The researchers set up a big test with 72 real clinical cases and created 1,147 specific checklists (called rubrics) to grade the AI's answers. These checklists were made by a special AI that had access to the doctors' actual decisions and medical books—information the testing AI didn't have. This "secret advantage" ensured the grading was fair and grounded in real medical reality.
The Main Discovery: More Robots Don't Always Mean a Better Brain
The team tried seven different ways to organize the AI. Some were simple single-brain models, while others were complex "agentic" systems where one main robot acted as a boss, hiring smaller robots to handle specific tasks like surgery, radiation, or drug therapy. Some of these robot teams even had a dedicated "fact-checker" robot to verify claims, and some could even hire new robots on the fly if they felt they needed extra help.
Here's the twist: Adding more tools and more robots didn't automatically make the AI smarter.
In fact, for some models, giving them access to the internet or PubMed (a medical database) actually made them worse at their job. It's like giving a student a library card and a calculator, but if they don't know how to ask the right questions or check their math, they might just get more confused. The study found that the "best" setup was a specific combination of a powerful model (Claude Opus 4.8) using a "Divide and Conquer" strategy with a fact-checker and the ability to spawn new helpers. Even then, this top-performing team only scored 0.594 ± 0.025 out of a perfect 1.0.
The "Good Enough" Problem
The paper suggests that while these AI systems can generate recommendations that look medically relevant, they are not yet good enough to work alone in a hospital. The best AI still missed the mark on a huge chunk of the criteria.
When a human oncologist (a cancer doctor) took a closer look at the top AI's answers, they found eight specific ways the AI kept failing:
- Wrong Justifications: The AI got the right answer but gave the wrong reason for it (like saying "take this pill because it's blue" when the real reason is "it kills cancer cells").
- Missing or Wrong Advice: It sometimes forgot to suggest a necessary treatment or suggested one that was dangerous.
- Fake Citations: It made up references or claimed a real paper said something it didn't.
- Overconfidence: The AI would sound 100% sure even when the medical evidence was shaky or the patient's story was missing key details.
- Outdated Info: It suggested treatments that are no longer the standard of care.
The human doctor spent about 1 hour and 35 minutes reviewing just one of these AI-generated plans. This highlights a major problem: if a doctor has to spend that long double-checking every single AI suggestion, it might be faster to just do the work themselves.
The "Stage" Surprise
The AI did better on early-stage cancer (Stage I) than on the trickier, more complex cases (Stage III). This makes sense because early-stage treatment often follows a strict, simple rulebook. But when the cancer gets more complex, the AI gets lost in the weeds, struggling to juggle surgery, radiation, and genetics all at once.
What the Paper Rules Out
The study explicitly argues against the idea that "more complex is always better." Just because you build a multi-agent system with fact-checkers and web search tools doesn't mean it will outperform a simple model. In some cases, the extra complexity introduced new errors. The paper also rules out the idea that these systems are ready for unsupervised clinical use; they are not a "plug-and-play" solution yet.
How Sure Are We?
The authors are very careful with their language. They suggest that agentic systems can be useful but remain insufficient for unsupervised use. They measured the scores using a rigorous scoring system and observed specific failure modes through human review. They did not claim to have "solved" breast cancer treatment with AI, nor did they claim the AI is better than a human doctor. Instead, they found that while the technology is promising, it still has significant gaps that need to be fixed before it can be trusted to make life-or-death decisions without a human looking over its shoulder.
In short, the AI is like a very smart intern who has read all the books but still needs a senior doctor to check their homework, because they sometimes make up facts, get confused by complex cases, and are too confident in their mistakes.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.