← Latest papers
💬 NLP

PanCanBench: A Comprehensive Benchmark for Evaluating Large Language Models in Pancreatic Oncology

PanCanBench is a comprehensive benchmark comprising 282 authentic pancreatic cancer patient questions and 3,130 expert-derived criteria that reveals significant variations in the factual accuracy and completeness of 22 large language models, highlighting that higher rubric scores and web-search integration do not necessarily guarantee clinical safety or reduced hallucinations.

Original authors: Yimin Zhao, Sheela R. Damle, Simone E. Dekker, Scott Geng, Karly Williams Silva, Jesse J Hubbard, Manuel F Fernandez, Fatima Zelada-Arenas, Alejandra Alvarez, Brianne Flores, Alexis Rodriguez, Stephen
Published 2026-03-03
📖 4 min read☕ Coffee break read

Original authors: Yimin Zhao, Sheela R. Damle, Simone E. Dekker, Scott Geng, Karly Williams Silva, Jesse J Hubbard, Manuel F Fernandez, Fatima Zelada-Arenas, Alejandra Alvarez, Brianne Flores, Alexis Rodriguez, Stephen Salerno, Carrie Wright, Zihao Wang, Pang Wei Koh, Jeffrey T. Leek

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a very smart, super-fast robot librarian who has read every medical book ever written. You ask it, "My uncle has pancreatic cancer; what should we do?" The robot gives you a long, confident answer.

The Big Problem: Just because the robot sounds confident and uses big words doesn't mean it's telling the truth. In the world of medicine, a confident lie can be dangerous.

This paper, called PanCanBench, is like a "final exam" designed specifically to test these robot librarians on the tricky, scary topic of pancreatic cancer. Here is how they did it, explained simply:

1. The "Real World" Test (No Fake Questions)

Most tests for AI use fake questions made up by computers (like "What is the capital of France?"). But real patients ask messy, emotional, and complex questions like, "My doctor said it's stage 4, but the biopsy wasn't clear. What now?"

The researchers went to a real help line for pancreatic cancer patients and grabbed 282 real questions from real people. They didn't just ask the AI to answer; they asked: "Is this answer safe and helpful?"

2. The "Human Grader" vs. The "Robot Judge"

To grade the answers, the researchers didn't just use a computer. They hired oncology fellows (doctors in training who specialize in cancer) to create a "grading rubric."

  • The Rubric is like a Recipe: Imagine a recipe for a perfect cake. It says: "Must have flour (5 points), must have sugar (5 points), must NOT have salt (-10 points)."
  • The doctors wrote these "recipes" for every single question.
  • Then, they used a super-smart AI (GPT-5) to act as the Judge, checking the other robots' answers against these human-written recipes.

The Result: The AI Judge was almost as good as the human doctors at grading. This means we can use AI to grade AI answers in the future, saving time and money!

3. The Shocking Findings

When they ran the test, here is what happened:

  • Confidence \neq Correctness: Some of the newest, "smartest" reasoning models (like o3) got the highest scores on the "recipe" test. They sounded perfect! BUT, they were also the ones making the most dangerous factual mistakes. It's like a student who writes a beautiful essay but gets the math wrong.
  • The "Hallucination" Rate: This is when the AI makes things up.
    • The best models (like GPT-5 and Gemini) only made up facts about 6% of the time.
    • Some open-source models (free models anyone can download) made up facts in 54% of their answers! That's like flipping a coin to decide if you should take your medicine.
  • The "Google Search" Trap: The researchers turned on the "Web Search" button for the robots, hoping they would look up the facts.
    • Surprise: It didn't always help! Sometimes, when the robot started searching the internet, it forgot the good information it already knew. It's like a chef who stops cooking to read a recipe book and burns the soup because they got distracted.
    • Also, the robots sometimes found bad links (like a random blog) instead of reliable medical journals.

4. The "AI vs. Human" Rubric Experiment

The researchers asked: "Can we just let an AI write the grading recipe (rubric) instead of hiring doctors?"

  • The Answer: No.
  • When AI wrote the grading recipes, the robots got much higher scores (inflated by nearly 18 points).
  • It's like if a student wrote their own test questions and then graded their own answers. They would get an A+, but they wouldn't actually know the material. Human experts are still needed to set the rules.

The Bottom Line

This paper is a warning label for the future of AI in healthcare.

  1. Don't trust the "smartest" sounding AI: The model that sounds the most confident isn't always the most accurate.
  2. Human experts are irreplaceable: We need doctors to help design the tests and check the facts.
  3. Open-source models need work: The free models are currently too risky for serious medical advice because they lie too often.

In short: PanCanBench is a reality check. It shows that while AI is getting better at talking like a doctor, it still needs a human supervisor to make sure it's not acting like a dangerous charlatan.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →