GAZE: Grounded Agentic Zero-shot Evaluation with Viewer-Level Tools and Literature Retrieval on Rare Brain MRI
The paper introduces GAZE, a grounded agentic framework that enables medical vision-language models to iteratively inspect brain MRIs using viewer-level tools and literature retrieval, significantly improving zero-shot performance on rare neurological conditions compared to standard single-pass models.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a radiologist looking at an MRI scan of a brain. They don't just glance at it once and guess what's wrong. Instead, they zoom in on a suspicious spot, adjust the brightness to see hidden details, flip the image to check for symmetry, and then open a medical textbook or search online to compare what they see with known cases. They do this iteratively, building a picture of the diagnosis step-by-step.
Current AI models, however, usually work like a student taking a speed-reading test: they get one image, think for a split second, and spit out an answer immediately. This works okay for common things, but when it comes to rare brain conditions, this "one-shot" approach often fails because the AI misses subtle clues or doesn't know enough about the specific disease.
The paper introduces GAZE (Grounded Agentic Zero-shot Evaluation), a new way to make AI behave more like that careful, iterative radiologist.
The "Smart Assistant" Analogy
Think of GAZE not as a single brain, but as a smart assistant that manages a team of tools for the AI.
The Viewer Tools (The Magnifying Glass):
Instead of just staring at the whole image, the AI can now use "viewer-level tools." It can:- Zoom in on a tiny spot.
- Adjust contrast (like turning up the brightness on a TV) to make dark spots pop.
- Detect edges to see the outline of a lesion clearly.
- Flip or rotate the image to check for symmetry.
- Reset the view if it gets lost.
- Analogy: It's like giving the AI a physical lightbox and a set of magnifying glasses, allowing it to inspect the image as a human would, rather than just processing a static picture.
The Library Tools (The Researcher):
If the AI sees something strange, it can pause and ask for help from two official medical libraries (PubMed for text and Open-i for images).- It searches for articles about the disease it suspects.
- It looks at pictures of that disease from other patients to see if they match.
- Analogy: It's like a detective who, after finding a clue, immediately runs to the library to check if that clue matches a known criminal profile before making an arrest.
How It Works (The "Agent" Part)
GAZE turns the AI into an agent. Instead of being forced to answer immediately, the AI is given a "continue" button.
- Step 1: The AI looks at the image. "Hmm, that looks weird."
- Step 2: It decides to zoom in on that spot.
- Step 3: It looks again. "Okay, it's definitely a lesion. But what kind?"
- Step 4: It decides to search the library for "brain lesions that look like this."
- Step 5: It reads the results, compares them, and finally writes its report.
The system records every single move the AI makes (every zoom, every search) so doctors can audit exactly how the AI reached its conclusion.
What They Found (The Results)
The researchers tested this on NOVA, a difficult dataset of 906 brain MRI scans featuring 281 different rare neurological conditions.
- The "Framework" Matters: Even before using any tools, just changing how the AI was asked to answer (using strict rules and better prompts) made it significantly better at finding lesions. It's like giving a student a better checklist; they do better even without new information.
- Tools Help the Rare Cases Most: The tools were a game-changer for rare diseases. For conditions that appeared very few times in the test set, the AI's ability to find the correct spot on the brain jumped from 17% to 58%. For common diseases, the improvement was good, but not as dramatic.
- The "Trade-Off" Surprise: The researchers found a tricky side effect. Sometimes, when the AI searched the library, it got the diagnosis right but moved the location of the lesion to the wrong spot.
- Analogy: Imagine the AI reads a book saying "This disease usually appears in the left ear." It then sees a spot in the right ear on the MRI. Instead of trusting its eyes, it moves its "guess" to the left ear because the book said so. It got the name of the disease right, but pointed to the wrong place. This proves you can't just test if an AI knows the disease name; you have to test if it can point to the right spot too.
- Engagement is Key: The best AI models (like Gemini 3 Flash) didn't just use the tools; they loved using them. They made about 12 tool calls per case, zooming and searching repeatedly. Weaker models barely used the tools at all, which is why they didn't improve much.
The Bottom Line
GAZE shows that to make AI good at medicine, especially for rare diseases, we can't just feed it a picture and ask for an answer. We need to build systems that let the AI look closer, search for answers, and think in steps, just like a human doctor.
The paper concludes that this "agentic" approach, combined with strict rules to ensure the AI doesn't get lost, significantly improves how well AI can find and describe rare brain conditions, though it warns that we must check both the diagnosis and the location to ensure the AI isn't getting tricked by its own research.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.