PathAgentBench: Benchmarking Evidence-Seeking Vision-Language Models on Whole-Slide Pathology Image
This paper introduces PathAgentBench, a comprehensive benchmark utilizing 1,822 whole-slide images and expert annotations to evaluate vision-language models on evidence-seeking capabilities in pathology, revealing that while models excel at reasoning over curated evidence, they significantly struggle with autonomously acquiring and localizing diagnostic regions directly from gigapixel images.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to solve a massive mystery hidden inside a city that is billions of times larger than a single grain of sand. This city is a "Whole-Slide Image" (WSI) of a tissue sample, a digital picture so huge it would take a human hours just to scroll through it. In the real world, a pathologist (a doctor who diagnoses diseases by looking at cells) doesn't just stare at one tiny spot. They start by looking at the whole city from a helicopter to find suspicious neighborhoods, then zoom in to check the streets, and finally zoom in all the way to inspect individual houses and people to find the culprit. This is called "evidence-seeking."
Recently, scientists have built super-smart computer brains called "Vision-Language Models" (VLMs). These are like AI detectives that can look at pictures and read text at the same time. They are getting really good at solving puzzles when someone hands them a specific photo and asks, "What's wrong here?" But there's a big question: Can these AI detectives actually find the right photo on their own in that giant city, or do they just need a human to point at the spot first? If an AI can solve a riddle but can't find the room where the riddle is hidden, it's not a very good detective yet. This is the gap this new research wants to measure.
The Great AI Detective Test: PathAgentBench
A team of researchers has built a new, super-challenging test called PathAgentBench to see if AI can really act like a human pathologist. Instead of just giving the AI a pre-selected picture and asking, "What is this?" (which is what most old tests do), they gave the AI the entire giant city and asked it to go hunting for the clues itself.
Think of it like a game of "Where's Waldo," but the book is the size of a football field, and you have to zoom in and out to find the tiny details that prove who the bad guy is. The researchers created a "diagnostic tree," which is like a map of a treasure hunt. The AI starts at the top (the whole slide), picks a suspicious area, zooms in, picks a smaller area, zooms in again, and finally makes a diagnosis based on all the clues it collected along the way.
To make sure the test was fair and tough, they used 1,822 real medical slides from a huge database and had ten expert human doctors draw the exact paths they would take to solve the cases. They then tested 20 different AI models (some made by big tech companies, some open-source, and some specialized for medicine) to see how they performed.
Here is what they found, and it's a bit of a plot twist:
1. The AI is a Genius at Reading, but a Terrible Navigator
When the researchers handed the AI the specific clues it needed (like, "Here is the zoomed-in picture of the suspicious cell, now tell me what it is"), the AI was amazing. The best models got over 93% of the answers right. They could read the evidence perfectly.
However, when the researchers said, "Okay, now you have to find that suspicious cell yourself in the giant city," the AI completely crashed. Even the smartest AI models achieved a spatial overlap score (called "mean intersection-over-union") of less than 0.09. This means the box the AI drew around the suspicious area barely touched the real spot. In fact, a very simple, dumb trick—just guessing the center of the previous box—actually worked better than the fancy AI models!
2. The "Zoom-Out" Problem
The researchers watched the AI try to explore the slides on its own. At the widest view (low magnification), the AI could find suspicious areas about 52% of the time. But as it tried to zoom in to get a closer look, it started losing the trail. By the time it reached the highest zoom level (where it needs to see the tiny details), it was only finding the right spot 2% of the time. It's like a detective who finds the right neighborhood but then gets lost in the wrong alleyway and never finds the house.
3. Specialized AI Isn't Always Better
You might think that AI models trained specifically on medical books would be the best detectives. Surprisingly, the general-purpose AI models (the ones trained on everything from cats to cars) often did better at finding the clues than the ones trained only on medicine. The specialized medical models sometimes got stuck or couldn't even output the coordinates to point at a spot.
4. The Cost of Hunting
The study also looked at how much "brain power" and money it took. The part where the AI had to hunt for the clues (the "autonomous exploration") was incredibly expensive and slow. It took the AI several minutes just to look at one slide, whereas answering a simple question about a pre-selected picture took a fraction of a second.
The Bottom Line
The main takeaway from this paper is that while AI has become incredibly good at interpreting evidence when it's handed to it, it is still very bad at finding that evidence on its own. The researchers argue that we shouldn't just celebrate AI that can solve riddles; we need to teach it how to be a detective who can actually find the riddle in the first place.
Until AI can reliably navigate these giant digital slides without getting lost, it can't fully replace a human pathologist in the real world. The paper suggests that future AI needs to learn better "spatial reasoning" (knowing where it is), how to use tools to zoom in and out correctly, and how to recover when it makes a wrong turn. For now, the AI is a brilliant reader, but it's still learning how to walk the streets.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.