NeuroQA: A Large-Scale Image-Grounded Benchmark for 3D Brain MRI Understanding
NeuroQA is a large-scale, image-grounded benchmark comprising nearly 57,000 question-answer pairs derived from 3D brain MRI volumes across 12 datasets and five clinical domains, designed to rigorously evaluate 3D medical visual reasoning while eliminating text-only shortcuts through strict image-grounding protocols and expert verification.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a robot to read a brain scan. The robot is incredibly smart; it has read millions of medical textbooks and knows the names of every part of the brain. But there's a catch: when you show it a 3D brain scan and ask, "Is there a tumor here?", the robot might not actually be looking at the scan. Instead, it might just be guessing based on the words in your question or the name of the patient group it was trained on. It's like a student who memorized the answer key to a test without ever studying the material.
This paper introduces NEUROQA, a new, massive "exam" designed specifically to stop robots from cheating and force them to actually look at the 3D brain images.
Here is how the paper breaks it down, using simple analogies:
1. The Problem: The "Text-Only" Cheat Code
Previous tests for medical AI were like giving a student a multiple-choice quiz where the questions were written in a way that gave away the answers.
- The 2D Problem: Most old tests showed the AI just a flat, 2D slice of a brain (like a single page of a book). But real brain scans are 3D volumes (like a whole book). You can't understand the whole story by reading just one page.
- The "Shortcut" Problem: The paper found that if you took the image away from these old tests, the AI still got 70–99% of the answers right! This means the AI wasn't "seeing" the brain; it was just guessing based on text patterns (e.g., "If the question mentions 'Parkinson's,' the answer is usually 'Yes'").
2. The Solution: NEUROQA (The "Honest" Exam)
The authors built a new benchmark called NEUROQA (56,953 questions from 12,977 real patients). Think of this as a rigorous, anti-cheating exam with three special rules:
- The 3D Rule: Every question comes with the entire 3D brain volume, not just a flat slice. The AI has to navigate the brain in 3D space, just like a doctor does.
- The "No-Image" Stress Test: To prove the AI is actually looking, they run a test where they hide the image. If the AI's score drops significantly when the image is gone, it proves the AI was actually using the image. If the score stays high, the AI is just guessing based on text.
- The "Human Ceiling": They didn't just compare the AI to random guessing; they compared it to real human doctors. Two doctors (a radiology resident and a neurosurgery resident) took the same test using a 3D viewer. Their average score was about 49%. This sets a "human visual limit." If an AI scores lower than a human looking at the same view, it hasn't mastered the task yet.
3. The Two Types of Questions
The exam has two types of questions, which is a crucial distinction:
- Image-Grounded (The "Look and See" Questions): These are things you can spot just by looking at the 3D scan, like "Is the left side of the brain smaller than the right?" or "What color is this spot on the scan?"
- Image-Informed (The "Math and Context" Questions): These require knowing specific numbers that you can't see with your eyes. For example, "Is the volume of this brain part below the 5th percentile?" You can't eyeball a precise milliliter measurement. The "correct" answer comes from a computer calculation (FreeSurfer), not just human vision. This tests if the AI can extract precise data, not just guess.
4. The Results: The AI is Still Struggling
The authors tested the smartest AI models available today (like the latest versions of GPT, Gemini, and Claude) on this exam.
- The Cheating Check: They cleaned up the questions so that if you removed the image, the AI would only get about 44.6% right (which is barely better than random guessing). This proves the exam is fair and doesn't give away answers.
- The AI Performance: Even the best AI models only scored around 47.5% on the closed questions (Yes/No and Multiple Choice).
- The Human Performance: The human doctors scored around 48.9%.
- The Conclusion: Currently, the best AI models are not beating the human doctors, and in some cases, they are scoring lower than the text-only baseline. This means the AI is not yet reliably "seeing" the 3D brain better than a human can, nor is it extracting the precise quantitative data needed for a diagnosis.
5. How They Built It (The "Factory")
To make sure the exam was perfect, they didn't use AI to generate the questions (which would be circular logic). Instead, they built a deterministic pipeline.
- Imagine a factory assembly line with 38 strict rules.
- They took real patient data and ran it through this machine.
- Human experts (doctors) reviewed the questions to make sure they made medical sense.
- They stripped away any "clues" in the text that might let the AI guess without looking.
- The result is a "gold standard" dataset where every answer is mathematically verified against the patient's actual scan data.
Summary
NEUROQA is a giant, honest test for medical AI. It forces the AI to look at full 3D brain scans and proves that, right now, even the smartest AI models are still struggling to understand these scans as well as a human doctor can. The paper doesn't claim the AI is ready for hospitals; instead, it provides a clear, measurable way to track when AI finally learns to "see" the brain properly.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.