← Latest papers
💻 computer science

Are Multimodal LLMs Ready for Clinical Dermatology? A Real-World Evaluation in Dermatology

This study shows that while multimodal large language models are promising on public dermatological benchmarks, their diagnostic accuracy declines significantly in real clinical settings, indicating that current models are not yet reliable enough for bedside use.

Original authors: Roy Jiang, Hyunjae Kim, Zhenyue Qin, Morten Lee, Margaret MacGibeny, Ailish Hanly, Angela Sadlowski, Shanin Chowdhury, Xuguang Ai, Jeffrey Gehlhausen, Qingyu Chen

Published 2026-05-07
📖 5 min read🧠 Deep dive

Original authors: Roy Jiang, Hyunjae Kim, Zhenyue Qin, Morten Lee, Margaret MacGibeny, Ailish Hanly, Angela Sadlowski, Shanin Chowdhury, Xuguang Ai, Jeffrey Gehlhausen, Qingyu Chen

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are hiring a new assistant to help a dermatologist (a skin doctor) sift through thousands of patient photos and notes. You want to know if this assistant is intelligent enough to look at an image of a skin rash, read the doctor's brief note about it, and say: "This looks like X, Y, or Z, and we need to see this patient immediately."

This article is like a rigorous "job interview" for five different AI assistants (so-called Multimodal Large Language Models or MLLMs) to test whether they are ready for this real-world job.

Here is the breakdown of what the researchers found using simple analogies:

1. The "Practice Exam" vs. The "Real Job"

The Setup:
Before hiring, the AI assistants took a "practice exam" using public datasets. These datasets are like perfectly staged photo shoots. The photos are high-quality, taken with specialized medical cameras (dermatoscopy), show only a clear skin area, and have no background clutter. In this perfect environment, the AI assistants performed quite well. The best one (GPT-4.1) scored about 42% correct answers, and the open-source models scored around 20–26%.

The Reality Check:
Then the researchers gave them the "Real Job" test. This included 5,811 actual patient cases from a hospital.

  • The Photos: These were not perfect studio shots. They were taken by regular doctors or patients with mobile phones. The lighting was poor, the image composition was messy, and there were often multiple rashes or distractions in the background (like a hospital bed or a hand holding the phone).
  • The Notes: The notes written by the referring doctors were often short, vague, or sometimes even gave a false impression of the problem.

The Result:
The performance of the AI assistants crashed.

  • When looking only at the messy real photos, the accuracy of the best AI model dropped from ~42% to ~24%.
  • The open-source models crashed even harder, dropping from ~26% to just 1.5% to 13%.
  • The Analogy: It is like a student who gets an "A" on a math exam with clean, printed numbers but fails miserably when asked to solve the same problems in illegible handwriting on a napkin.

2. The Trap of "Trusting the Note"

The researchers tested what happens when they give the AI the doctor's written note along with the photo.

  • The Good News: When the notes were helpful, the AI got better. It is like giving a detective a helpful clue.
  • The Bad News: The AI was too trusting. If the referring doctor wrote a note saying "I think this is a spider bite," the AI often ignored the image and simply agreed with the note, even when the image clearly showed something else.
  • The Danger: If the referring doctor was wrong (which happens often in real life), the AI would confidently repeat the wrong diagnosis. In some cases, a misleading note made the AI's performance worse than if it had only looked at the photo. One model (SkinGPT-4) was confused in 84% of cases when the note was misleading.

3. The "Triage" Test (Who Needs Help Now?)

Since the AI struggled to name the exact disease, the researchers asked a simpler question: "Is this patient's condition severe and urgent?" (Like a triage nurse deciding who goes next).

  • The Result: The AI was okay at this, but not great. It could identify about 60% of severe cases when notes were provided.
  • The Catch: It missed many severe cases, and its ability to recognize them depended entirely on the quality of the notes. When the notes were confusing, the AI missed the urgent cases.

4. Where Did They Fail? (The "Eyes" vs. The "Brain")

The researchers asked human dermatologists to evaluate the AI's reasoning. They found a surprising truth:

  • The AI's "brain" (the reasoning) was fine. The AI knew medical facts and could formulate sentences that sounded professional.
  • The AI's "eyes" (the image recognition) were the problem. The AI failed because it could not accurately describe the specific visual details of the rash. It said general things like "there is a red spot," instead of noticing the specific shape or texture that defines a disease.
  • The Analogy: Imagine a student who has memorized the entire dictionary and can write a perfect essay, but when shown a picture of a cat, cannot tell you that it has whiskers or pointed ears. They know the words for a cat, but they cannot see the cat.

The Bottom Line

The study concludes that these AI models are not ready to work alone in a real dermatological practice.

  • Benchmark performance is misleading: Just because an AI performs well on clean, perfect datasets does not mean it can handle the messy reality of a hospital.
  • The bottleneck is not "thinking": The problem is not that the AI cannot reason; it is that it cannot accurately "see" the skin in real photos and is easily confused by poor notes.
  • Safety Warning: As long as the AI does not get better at looking at messy photos and ignoring misleading notes, it cannot be trusted to make decisions without a human doctor looking over its shoulder.

In short: The AI is a smart student who passed the written exam but failed the practical driving test. It needs more training on real "road conditions" before it can drive the car.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →