ScintiGround-38K: A Grounded Visual Question Answering Benchmark for Paired Anterior–Posterior Bone Scintigraphy
The paper introduces ScintiGround-38K, a large-scale, provenance-preserving benchmark for grounded visual question answering on paired anterior-posterior bone scintigraphy images, which reveals that while models achieve high raw answer accuracy, they struggle significantly with minority-class detection and strict evidence-based localization.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to solve a mystery, but instead of a crime scene, you are looking at a glowing map of a human skeleton. This map, called a bone scan, is created by injecting a tiny amount of safe, radioactive dye into a patient. The dye travels through the blood and sticks to areas where bones are growing or healing, lighting them up like neon signs on a dark road. Doctors use these scans to spot trouble, like cancer that has spread to the bones, by looking for these bright "hotspots."
Now, imagine you want to teach a computer to be a detective too. You give it a picture of the skeleton and ask, "Is there a problem here?" This is called Visual Question Answering (VQA). But here is the tricky part: a computer can sometimes be a "smart guesser." It might look at the question, guess the answer based on how often it sees certain words, and say "Yes" without actually looking at the glowing spots on the map. It's like a student who memorizes the answer key but doesn't understand the math. To fix this, scientists are building a new kind of test where the computer must not only give the right answer but also point its finger at the exact spot on the picture that proves it. This is called "grounded" answering. The big question is: can these smart computers actually learn to look at the evidence, or are they just guessing?
This is exactly what the researchers behind ScintiGround-38K set out to find. They created a massive, super-organized game for computers to play using 38,469 different bone scan puzzles. They didn't just make up questions; they built the game using real medical records from 2,815 patients, keeping the original "glowing" spots and the exact coordinates where doctors had marked them. They turned these records into three types of challenges: asking a yes-or-no question, asking the computer to draw a box around a specific bone, and asking it to do both at the same time.
The results were a bit of a shocker, like finding out a champion chess player is actually just guessing the moves. When the researchers tested two of the smartest AI models available (Qwen3-VL and PaliGemma), the computers got the simple "Yes or No" answers right more than 90% of the time. That sounds amazing, right? But when the researchers looked closer, they found a hidden trap. Because most of the bone scans in the game were perfectly normal, the computers learned to just guess "No problem" almost every time. When they forced the computers to prove their answer by pointing to the glowing spot, the performance dropped significantly.
The models were actually quite good at drawing boxes around the glowing spots (getting about 70-85% of the boxes right), but they were terrible at spotting the rare, abnormal spots. Their ability to correctly identify the "bad" spots was so low that it was barely better than flipping a coin. The paper shows that these computers are currently "answer experts" but "evidence novices." They can say the right words, but they often fail to connect those words to the actual glowing evidence on the map.
The authors are very clear about what this means: this is not a tool ready to diagnose patients in a hospital yet. It is a research tool designed to show us exactly where current AI is failing. It proves that getting a high score on a simple test is misleading if the computer isn't actually looking at the picture. The paper concludes that for AI to be truly trustworthy in medicine, we need benchmarks like ScintiGround-38K that force the computer to show its work, ensuring it isn't just guessing based on patterns in the questions, but is truly seeing the evidence. Until we can fix the "guessing" problem, these smart systems remain powerful research tools, but not yet ready for the real world of patient care.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.