← Latest papers
💻 computer science

Comparative Performance of Frontier Large Language Models for Extracting High-Risk Pathologic Features from Unstructured Gastrointestinal Oncology Reports: A Systematic Benchmarking Study with Human and Traditional NLP Baselines

This systematic benchmarking study demonstrates that frontier large multimodal models, particularly Gemini 2.5 Pro and GPT-4o, significantly outperform both time-pressured human experts and traditional NLP baselines in extracting critical perineural invasion status from unstructured gastrointestinal pathology reports, while a novel failure mode taxonomy provides actionable guidance for model selection based on specific error signatures.

Original authors: Seher Siddiqui

Published 2026-07-21✓ Author reviewed
📖 4 min read☕ Coffee break read

Original authors: Seher Siddiqui

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the world of medical records as a giant, chaotic library where every book is written in a different dialect. In this library, doctors write long, messy stories about what they see under a microscope, and hidden inside those stories are tiny, crucial clues that decide a patient's future treatment. One of the most important clues is something called "perineural invasion" (PNI). Think of PNI as a sneaky thief that lets cancer cells sneak along the tiny nerves in the body, making the disease harder to fight. If a doctor misses this clue, they might not give the patient the right kind of radiation or extra medicine they need.

For years, computers have tried to read these stories to find the clues, but they often get confused by the messy handwriting of human language. They might miss a "no" or get stuck on a complicated sentence. Now, a new generation of super-smart computer brains, called "Large Language Models" (like the ones that power chatbots), has arrived. These aren't just simple search engines; they are like brilliant students who can read a whole story, understand the context, and figure out what the doctor really meant, even if the words are tricky. The big question everyone is asking is: Are these new AI students smart enough to find the sneaky cancer clues better than a tired human doctor or an old-school computer program?

This paper is like a giant, high-stakes spelling bee, but instead of spelling words, the contestants are trying to find those sneaky cancer clues in 445 real medical reports. The researchers set up a race between four of the most powerful AI models available today—GPT-4o, Gemini 2.5 Pro, Claude 3.7 Sonnet, and DeepSeek-R1. But they didn't just let the AI play alone; they added two other teams to the race to see who would win. The first team was a group of human doctors-in-training (residents) who had to read the reports in a frantic 90-second time limit, simulating how busy real doctors are. The second team was a traditional, older computer program (BioBERT) that had been specifically trained on medical text but lacks the "common sense" of the newer AI.

The results were a clear victory for the new AI models, but with some interesting twists. The winner of the race for pure accuracy was Gemini 2.5 Pro, which got the right answer 95.8% of the time. It was followed very closely by GPT-4o, which got 94.2% right. Both of these AI superstars significantly outperformed the human residents, who only got 77.2% right under the time pressure, and the old-school BioBERT program, which scored 79.6%. In fact, the best AI model was nearly 10% more accurate than the rushed human team.

However, the paper also found that the AI isn't perfect yet. Even the best AI models still didn't do as well as a team of expert pathologists who took their time and didn't have a stopwatch ticking over their heads (who got 97.2% right). This means the AI is great at doing the first pass of the work, flagging the important cases, but it still needs a human expert to double-check the final decision.

The study also discovered that each AI model makes different kinds of mistakes, like having different "personality flaws." Gemini 2.5 Pro was the best at getting the overall score right, but it sometimes got confused by negative words. If a report said "no evidence of invasion," Gemini sometimes missed the "no" and thought the cancer was there. On the other hand, GPT-4o and Claude 3.7 Sonnet were very careful. If the report used vague words like "possible" or "likely," these models would often say, "I'm not sure," rather than guessing. While this is safe, it means they might miss some cases that a human would catch.

In the end, the paper suggests that these new AI models are ready to be used as a "first line of defense" in hospitals. They can scan through hundreds of reports much faster and more accurately than a tired human can, finding the dangerous clues that need attention. But because they still make specific types of errors and aren't quite as perfect as a careful expert, the best plan is to use the AI to do the heavy lifting and then have a human doctor give the final thumbs-up. It's a team effort where the AI is the super-fast scout, and the human is the wise captain making the final call.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →