← Latest papers
🤖 machine learning

NeuroVLM-Bench: Evaluation of Vision-Enabled Large Language Models for Clinical Reasoning in Neurological Disorders

This paper introduces NeuroVLM-Bench, a comprehensive benchmark evaluating twenty frontier vision-enabled large language models on 2D neuroimaging tasks, revealing that while technical image attributes are well-handled, diagnostic reasoning remains challenging, with proprietary models like Gemini-2.5-Pro leading in accuracy and MedGemma-1.5-4B showing the most promise among open-weight architectures.

Original authors: Katarina Trojachanec Dineva, Stefan Andonov, Ilinka Ivanoska, Ivan Kitanovski, Sasho Gramatikov, Tamara Kostova, Monika Simjanoska Misheva, Kostadin Mishev

Published 2026-03-27
📖 6 min read🧠 Deep dive

Original authors: Katarina Trojachanec Dineva, Stefan Andonov, Ilinka Ivanoska, Ivan Kitanovski, Sasho Gramatikov, Tamara Kostova, Monika Simjanoska Misheva, Kostadin Mishev

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a team of 20 brand-new, super-smart robots. These robots are like "digital doctors" who can look at pictures of human brains (MRI and CT scans) and try to figure out what's wrong with the patient. Some of these robots are free and open-source (like a community-built tool), while others are expensive, proprietary "black boxes" owned by giant tech companies.

The big question is: Can these robots actually help real doctors, or are they just fancy guessers that might make dangerous mistakes?

This paper, titled NeuroVLM-Bench, is like a massive, rigorous "driver's license test" for these AI robots, specifically designed for brain disorders like tumors, strokes, and Multiple Sclerosis.

Here is the story of how they tested them, what they found, and what it means for the future.

1. The Exam: Not Just a Multiple-Choice Quiz

Most AI tests are like simple trivia games: "Show me a picture of a brain, and tell me if it's sick or healthy."

But the authors of this paper said, "That's not how real doctors work." Real doctors don't just say "Sick." They write detailed reports. So, they built a test where the AI had to fill out a structured medical report for every single brain scan. The AI had to correctly identify:

  • The Modality: Is this an MRI or a CT scan? (Like asking, "Is this a photo or a video?")
  • The Angle: Is the picture taken from the top, side, or front?
  • The Sequence: What specific settings were used for the MRI?
  • The Diagnosis: What is the disease? (e.g., Stroke, Tumor, Normal)
  • The Subtype: What kind of tumor? (e.g., Glioma vs. Meningioma)

If the AI got the diagnosis right but messed up the report format (like writing a paragraph instead of a checklist), it failed the test. This is crucial because real hospital computers need data in a specific format to work.

2. The Training Camp: A Three-Stage Filter

They didn't just throw all 20 robots into the final exam. They used a "survival of the fittest" approach:

  • Phase 1 (The Screening): They tested all 20 robots on a small practice exam. Many robots failed miserably. Some refused to answer (too scared), and some just made up nonsense. They cut the list down to the top 11.
  • Phase 2 (The Stress Test): The remaining 11 took a bigger, harder test to see if they could stay consistent. They cut the list down to the top 6.
  • Phase 3 (The Final Exam): The final 6 took the ultimate test on brand-new data they had never seen before. They took this test twice: once with no help (Zero-Shot) and once with a "cheat sheet" of 4 example cases to study first (Few-Shot).

3. The Results: The Good, The Bad, and The "Almost There"

The Winners (The "A-Students")

  • Gemini 2.5 Pro and GPT-5 Chat were the top performers. They were the most accurate at diagnosing complex brain issues.
  • Gemini 2.5 Flash was the "smartest value." It was almost as good as the top models but much faster and cheaper to run. Think of it as a sports car that gets great gas mileage.

The "Specialists" vs. The "Generalists"

  • Tumors: The robots were surprisingly good at spotting brain tumors. It's like they have a "sixth sense" for big, obvious lumps.
  • Strokes: They were okay at spotting strokes, but it was a bit hit-or-miss.
  • Multiple Sclerosis (MS): This was the hardest. The robots struggled significantly. MS lesions are tiny and scattered, like finding specific grains of sand on a beach. The robots often missed them or got confused.
  • Rare Diseases: When the brain had a weird, rare infection that looked like a tumor, the robots often got it wrong. This is dangerous because treating a tumor (surgery) is very different from treating an infection (antibiotics).

The "Cheat Sheet" Effect (Few-Shot Prompting)

When the researchers gave the robots a few examples to study before the test, some got much better.

  • MedGemma 1.5 4B: This is a smaller, open-source model (free to use). With the cheat sheet, it jumped up to perform almost as well as the expensive, giant models. This is huge news because it means we might not need billion-dollar models to get good results; we just need to teach them better.
  • The Catch: Using the cheat sheet made the robots slower and more expensive to run because they had to read more text.

4. The Big Problem: "Confident Wrongness"

The paper found a scary pattern. Some robots were very confident in their answers even when they were wrong.

  • The "Hallucination" Risk: Sometimes, if you erased a tumor from a picture, the robot would still say, "Yes, there is a tumor here." It was making things up based on what it thought it should see, not what was actually there.
  • Calibration: The best robots knew when they didn't know. They would say, "I'm not sure," or "I can't tell." The worst ones would guess wildly and pretend they were 100% sure. In a hospital, a robot that admits uncertainty is safer than a robot that is confidently wrong.

5. The Bottom Line: What Does This Mean for You?

Don't expect a robot to replace your neurologist tomorrow.
These AI models are like very bright medical students who are great at memorizing facts and spotting obvious patterns (like big tumors) but still struggle with the subtle, complex reasoning required for difficult cases like MS or rare infections.

The Future is "Teamwork," not "Replacement."
The paper suggests that in the future, hospitals won't just use one AI. They will use a smart routing system:

  • If a patient has a likely tumor, send the scan to the "Tumor Specialist AI."
  • If it looks like a stroke, send it to the "Stroke AI."
  • If the AI is confused or the case is rare, it should immediately flag it for a human doctor to review.

The Takeaway:
We are making incredible progress. The technology can already read the "metadata" (what kind of scan it is) perfectly. But the "medical reasoning" (figuring out the complex disease) is still a work in progress. The goal now isn't just to make the AI smarter, but to make it safer, cheaper, and better at knowing when to ask for human help.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →