← Latest papers
📄 medicine

Performance and Failure Patterns of Multimodal Large Language Models for Hepatocellular Carcinoma Detection on Portal Venous Phase CT

While the best-performing multimodal large language model (ChatGPT-5.1 with anatomical guidance) showed improved accuracy for detecting hepatocellular carcinoma on portal venous phase CT, it still fell significantly short of human radiologists, highlighting the current limitations of AI in visual lesion detection and the critical need for meticulous prompt design and human supervision.

Original authors: Liu Jiang, Fangshi Lu, Haibin Zhu, Anna S Samuel, Ciara O'Brien, Isabelle Danika Gauthier, Chelsea CY Kim, Xiaoting Li, Masoom Abbas Haider, Xiaoyang Liu

Published 2026-08-10
📖 4 min read☕ Coffee break read

Original authors: Liu Jiang, Fangshi Lu, Haibin Zhu, Anna S Samuel, Ciara O'Brien, Isabelle Danika Gauthier, Chelsea CY Kim, Xiaoting Li, Masoom Abbas Haider, Xiaoyang Liu

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a super-smart robot how to spot a hidden treasure in a giant, foggy forest. In the world of medicine, that "forest" is the human body, and the "treasure" is a dangerous tumor called hepatocellular carcinoma (HCC). Usually, doctors use a special kind of X-ray called a CT scan to see through the fog. But recently, a new type of robot brain has appeared: the Multimodal Large Language Model (LLM). Think of these models as incredibly well-read students who have studied millions of books and pictures. They are great at answering questions and chatting, but can they actually see a tumor in a blurry medical image the way a human doctor can? This is the big question. Doctors care because if these robots can learn to spot cancer early, they could help save lives. But if they get confused by the fog and miss the treasure, or if they think a harmless rock is a treasure, that could be dangerous. So, scientists wanted to put these AI students to the test to see if they are ready for the real world or if they still need a lot of studying.

In this study, a team of researchers from universities in China and Canada decided to play a high-stakes game of "Spot the Difference" with these AI models. They gathered 211 single-slice CT images of livers—some with confirmed liver cancer and some perfectly healthy ones. They handed these images to four different versions of AI "students" (including ChatGPT-4o, ChatGPT-4.5, Perplexity, and the newest ChatGPT-5.1) and asked them a simple question: "Is there a tumor here?" To make it fair, they also asked two human experts—a senior radiologist (a master detective with years of experience) and a fellow (a trainee detective)—to look at the exact same pictures and answer the same question.

The results were a mix of impressive progress and clear reminders of how much work is left to do. The human experts were the undisputed champions. The senior radiologist caught every single tumor (100% sensitivity) and was correct 91.5% of the time overall. The trainee fellow was right next to them, getting it right 91.0% of the time. The AI models, however, struggled a bit more. The best-performing AI, ChatGPT-5.1, managed to get 83.4% of the answers right. That's a solid score, but it was still significantly lower than the human doctors.

Here is where the story gets interesting: the researchers discovered that the AI's performance depended heavily on how they were asked the question. When they gave the AI just a standard prompt, it got about 74% right. But, when they added a tiny bit of "anatomical guidance"—basically telling the AI, "Hey, remember, the liver is on the left side of this image, and the lungs are behind it"—the AI's score jumped up to 83.4%. It was like giving the robot a flashlight and a map; suddenly, it could see much better. Even with this boost, though, the AI still missed more tumors than the human doctors did.

The study also looked closely at why the AI and the humans made mistakes. The AI's biggest failure was missing tumors that looked exactly like the healthy liver tissue around them (called "isoattenuating" lesions). It was like trying to find a white snowball in a pile of white snow. The humans were much better at spotting these tricky ones. On the flip side, both the AI and the humans sometimes got confused by blood vessels, mistaking a cross-section of a vein for a tumor. Interestingly, the AI made fewer of these "false alarm" mistakes than the humans did, but it had its own weird errors, like getting too obsessed with specific shades of gray and ignoring the bigger picture.

Ultimately, the paper suggests that while these AI models are getting smarter and can be helped by better instructions, they are not yet ready to replace human doctors. They are like very talented interns who are still learning the ropes. They can be useful tools to help doctors, but they need a human supervisor to double-check their work, especially when the clues are subtle. The researchers conclude that we need to be careful, keep testing these models, and understand exactly where they might fail before we trust them with real patient diagnoses.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →