Evaluation of Medical Vision Language Models HuluMed and MedGemma, and general purpose chatbots Gemma 3, ChatGPT Plus, and Claude Pro on real previously unseen wound images
This study evaluates six vision-language models on real wound images and finds that general-purpose systems like ChatGPT and Claude significantly outperform medically specialized models in clinical wound assessment, though all models still face substantial limitations in autonomous reliability.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a group of six different "digital detectives." Their job is to look at photos of 20 different types of open wounds (like cuts, sores, or surgical sites) and answer 12 specific questions about them, such as: "Is this infected?", "Do we need surgery?", or "What kind of bandage should we use?"
The researchers wanted to see which detective was the best at solving these medical puzzles. They pitted two types of detectives against each other:
- The Generalists: Famous, powerful AI chatbots (like ChatGPT and Claude) that know a little bit about everything, including medicine.
- The Specialists: AI models built specifically for doctors (like HuluMed and MedGemma) that were trained only on medical textbooks and data.
Here is what happened when they took the test:
The Race Results
The Generalists won the race by a landslide.
- ChatGPT was the top scorer, getting about 73% of the answers right.
- Claude came in second with about 62%.
The Specialists struggled significantly.
- HuluMed (the best of the specialist group) only got 40% right.
- The other medical AIs (MedGemma and Gemma 3) scored even lower, with one of them getting less than 18% correct.
The "Why" Behind the Scores
The paper explains this surprising result with a few key ideas:
1. The "Jack of All Trades" vs. The "Narrow Expert"
Think of the Generalist AIs as a Swiss Army Knife. They have seen millions of different images and conversations. When they look at a wound, they use their massive, broad experience to figure out the context.
The Specialist AIs are like a screwdriver designed only for one specific type of screw. While they know a lot about medical words, they seem to get confused when they have to look at a messy, real-world photo and decide what to do about it. The paper suggests that being trained on a huge variety of data helps these models understand the "big picture" better than being trained only on medical data.
2. Seeing vs. Deciding
The detectives were good at seeing things. Most of them could correctly identify if a wound looked red (infected) or had dead tissue (necrosis).
However, they were terrible at deciding. When asked, "Should we cut this out?" or "Do we need an MRI?", the Specialist AIs often gave vague, hesitant answers or suggested the wrong procedure.
- Analogy: It's like a student who can perfectly describe a car engine but fails when asked to actually drive the car to the destination.
3. The "Waffling" Problem
The researchers gave the AIs a strict grading rule: You must give a clear "Yes" or "No" answer. If you say, "It might be infected, but maybe not," you get zero points.
The Specialist AIs loved to "waffle." They would say things like, "It is possible that an intervention is needed, depending on the clinical signs." While this sounds careful and safe, the test marked it as wrong because it wasn't a clear decision. The Generalist AIs were more direct and confident in their answers.
The Big Takeaway
The paper concludes that right now, if you need an AI to look at a wound photo and help a doctor make a plan, the general-purpose chatbots are currently better than the medical-specific ones.
However, there is a big warning label attached to this result. The paper says these AI tools are not ready to work alone. They are like a very smart intern who can help a doctor organize their thoughts, but they still make mistakes. Doctors must always double-check the AI's work before making any real medical decisions.
In short: The "general knowledge" AIs are currently better at solving wound puzzles than the "medical school" AIs, but neither is perfect enough to replace a human doctor yet.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.