← Latest papers
💬 NLP

Can Language Models Identify Side Effects of Breast Cancer Radiation Treatments?

This study evaluates seven large language models in identifying breast cancer radiation side effects using a stress-testing framework, revealing their sensitivity to documentation changes and tendency to under-report rare or long-term toxicities, while demonstrating that grounding outputs in clinician-curated lists significantly improves reliability for survivorship care applications.

Original authors: Natalie Seah, Danielle S. Bitterman, Daphna Spiegel, Thomas Hartvigsen

Published 2026-05-12
📖 5 min read🧠 Deep dive

Original authors: Natalie Seah, Danielle S. Bitterman, Daphna Spiegel, Thomas Hartvigsen

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a patient who just finished radiation treatment for breast cancer. You are about to leave the hospital, and your doctor needs to give you a list of things to watch out for in the coming weeks and years. This list is crucial: it helps you know what's normal, what needs attention, and what might happen years down the road.

Now, imagine asking a super-smart robot (a Large Language Model, or LLM) to write that list for you. The paper you are reading asks: Can we trust this robot to get the list right?

Here is the story of what the researchers found, explained simply.

The Setup: A Stress Test for Robots

The researchers didn't just ask the robots to "guess" side effects. They set up a "stress test," like a crash test for cars.

  1. The Patients: They created 21 fake patient profiles. These were like detailed character sheets for video game characters, containing age, medical history, and treatment details.
  2. The Twist: For every single patient, they created two versions of the story.
    • Version A (Vague): "The patient had radiation."
    • Version B (Specific): "The patient had radiation on the left breast and chest wall."
    • Everything else was exactly the same.
  3. The Goal: They asked seven different AI models to list the side effects for both versions. They then compared the AI's lists to a "Gold Standard" list created by a team of real breast cancer doctors.

The Findings: Where the Robots Stumbled

1. The "Guessing Game" Problem (Free-Form Mode)

When the researchers let the AI write the list however it wanted (like asking a friend to "tell me everything"), the results were a mixed bag.

  • The "Safe" Robots: Some models were very careful. They only listed side effects they were 100% sure about. They rarely made mistakes (high precision), but they missed a lot of important things (low recall). It was like a librarian who only recommends books they have read cover-to-cover, missing many great stories.
  • The "Chatty" Robots: Other models tried to be helpful by listing everything they could think of. They caught more of the right side effects, but they also made up a lot of nonsense. They listed side effects for radiation on the pelvis when the patient only had radiation on the chest. This is like a tour guide who knows the city well but keeps pointing out landmarks in the wrong neighborhood.

2. The "Counting" Trap

The researchers tried to fix the "Chatty" robots by saying, "Hey, just give us a list of 20 to 30 side effects."

  • The Result: This backfired. When the robots were forced to hit a specific number, they started making up fake side effects just to fill the quota. It's like asking a child to write 20 sentences about their day; eventually, they start inventing things just to reach the count. The lists became less accurate, not more.

3. The "Tiny Change" Sensitivity

This was one of the most surprising findings. The robots were incredibly sensitive to tiny changes in the text.

  • If the text said "radiation," the robot might list 10 side effects.
  • If the text said "radiation on the left breast," the robot might suddenly list 15 different side effects, or forget the first 10.
  • The Metaphor: Imagine a weather app that predicts "sunny" if you type "New York," but predicts "snow" if you type "New York, NY." The output shouldn't change so drastically just because you added two letters. The robots were unstable; they couldn't agree with themselves when the details were slightly tweaked.

4. The "Long-Term" Blind Spot

The researchers looked at when side effects happen.

  • Short-term: Things like skin redness or tiredness.
  • Long-term: Things like heart issues or scarring that might appear years later.
  • The Finding: Most robots were great at listing the short-term stuff but terrible at remembering the long-term stuff. They were like a news anchor who only reports on the breaking news of the day but forgets to mention the historical context. This is dangerous for cancer survivors who need to know about risks that might show up a decade later.

The Solution: The "Menu" Approach

The researchers found one way to make the robots much more reliable: Give them a menu.

Instead of asking the robot to "invent" a list, they gave the robot a pre-written list of side effects created by real doctors and asked, "Which of these apply to this patient?"

  • The Result: The robots became much smarter. They stopped making up fake side effects (no more hallucinations) and became much better at picking the right ones.
  • The Metaphor: It's the difference between asking a chef to "make a meal" (where they might grab random ingredients from the wrong aisle) versus giving them a specific menu of ingredients and asking them to "pick the right ones for this recipe." The second way is much safer and more consistent.

The Bottom Line

The paper concludes that while AI can be a helpful tool, we cannot just let it "freely write" medical advice.

  • If you let it guess, it might miss important long-term risks or invent scary things that aren't real.
  • If you force it to count to a specific number, it will lie to fill the space.
  • However, if you give it a trusted, doctor-approved list to choose from, it becomes a much more reliable assistant.

The researchers are essentially saying: Don't let the AI be the author of the medical advice; let it be the editor who checks a list written by experts. This ensures that cancer survivors get accurate, safe, and complete information about their future health.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →