PathoArgus: Advancing Evidence-Grounded Long-Context Visual Reasoning across Gigapixel Whole-Slide and Multi-Slide Case Contexts
This paper introduces PathoArgus-Bench, a comprehensive benchmark and evaluation protocol that exposes the critical gap between answer accuracy and true evidence-grounded reasoning in computational pathology by demonstrating that even high-performing models fail to consistently link gigapixel whole-slide image evidence to their predictions.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the quiet, sterile world of a pathology laboratory, a diagnosis often begins with a single, massive image. This is a whole-slide image, a digital photograph of a tissue sample so large and detailed that it contains billions of tiny pixels, far more than a human eye could ever scan in a single glance. To make sense of this, a pathologist must act like a detective, searching through vast landscapes of cells to find the specific, tiny clues that confirm a disease. For years, scientists have been teaching computers to do this same work, hoping to build artificial intelligence that can read these gigapixel images and answer questions about a patient's health. The goal has always been to create a system that doesn't just guess the right answer, but actually finds the evidence in the tissue to prove it. However, a critical question has remained unanswered: when a computer gets the right answer, is it because it truly saw the disease in the slide, or is it simply guessing based on the way the question was asked?
A new study from researchers at the Hong Kong University of Science and Technology tackles this problem head-on by introducing a rigorous new way to test these computer systems. They built a massive testing ground called PathoArgus-Bench, which contains over 22,000 questions derived from real patient cases across 15 different types of cancer. Unlike previous tests that only checked if the final answer was correct, this new benchmark forces the computer to prove it looked at the right place. The researchers created a strict rule: the computer is only allowed to look at a tiny fraction of the massive image, simulating the real-world limit of how much data a system can process at once. They then asked the computers to answer questions based on this limited view. To see if the computers were truly reasoning, the researchers also created a special set of tests where they swapped the tissue samples around. If a computer was truly looking at the evidence, changing the tissue should change its answer. If it was just guessing, it would likely give the same answer regardless of what was actually in the picture.
The results of this experiment revealed a startling gap between what computers appear to know and what they actually understand. When the researchers tested 20 different artificial intelligence systems, including the most advanced models available, they found that while some could get the right answer more than half the time, they were almost entirely failing to follow the evidence. One of the top-performing systems, a powerful model known as GPT-5.6, achieved a 57 percent success rate on the questions. However, when the researchers checked if the system changed its mind when the tissue evidence was altered, it succeeded in doing so only about 4 percent of the time. In other words, the system was often getting the right answer for the wrong reasons, relying on patterns in the text of the question rather than the visual reality of the tissue. Even when the researchers built a new tool specifically designed to help the computer find the most relevant parts of the image, the system still struggled to consistently link its answer to the specific evidence it found.
This discovery suggests that the current generation of medical artificial intelligence is not yet ready to be trusted with the complex task of diagnosing disease from whole-slide images. The study shows that simply giving a computer access to a large image is not enough; the system must also be able to navigate that image, find the specific clues, and let those clues dictate its conclusion. The researchers found that while acquiring useful whole-slide context is necessary, it is far from sufficient, and that current systems exhibit a stark gap in evidence-grounded reasoning. This indicates that the ability to find evidence and the ability to reason based on that evidence are two separate skills, and current systems have not yet mastered the latter. The work serves as a clear warning that measuring success by the final answer alone is misleading. To truly advance the field, the focus must shift from simply getting the right score to ensuring that the computer's reasoning is grounded in the actual visual facts of the patient's tissue.
The study concludes that while we have made progress in teaching computers to see, we have not yet taught them to think with what they see. The new benchmark and the tools developed in this research provide a roadmap for the next generation of medical AI, one that demands proof of evidence rather than just a correct guess. By exposing the limitations of current systems, the researchers hope to guide future development toward models that can truly understand the complex visual language of disease, ensuring that when a computer makes a diagnosis, it is because it has seen the truth, not because it has learned a trick.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.