Polish Medical Visual Question Answering: Vision-Language Models Underutilize Visual Evidence
This paper introduces a Polish medical visual question answering benchmark derived from specialist certification exams and reveals that current vision-language models underutilize visual evidence, often relying more on text cues and answer choices than on the actual images to solve the tasks.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a robot how to be a doctor. You wouldn't just give it a stack of textbooks; you'd also need to show it X-rays, heart monitor squiggles, and photos of rashes. This is the world of Vision-Language Models (VLMs). Think of these models as super-smart students who can read a medical textbook and look at a picture of a broken bone at the same time. They are supposed to combine what they see with what they know to answer tricky questions. But here's the big worry: what if the robot is just a "cheater"? What if it ignores the picture entirely and guesses the answer just by reading the multiple-choice options or the text description, like a student who memorized the test answers without actually studying the subject? This paper dives into that exact question, asking if our AI doctors are truly looking at the evidence or just skimming the clues.
The researchers behind this study decided to put Polish medical AI to the ultimate test. They built a special exam using real questions from the Polish Board Certification Examination, the tough test that real doctors and dentists must pass to become specialists. They took 286 of these questions that included images—like CT scans, heart wave charts, and clinical photos—and asked various AI models to solve them. They also created a control group of text-only questions to see how the models did without any pictures at all.
Here is what they found, and it's a bit of a plot twist. The best AI model managed to get about 79.0% of the answers right on the full image-based test. That sounds impressive, but when they compared it to actual human doctors taking the same exam, the humans scored around 70.14%. So, one specific commercial AI (GPT-5.6) actually beat the humans, but almost every other model they tested performed worse than a human doctor would.
But the real magic happened when the researchers started playing "hide and seek" with the information. They tested the models in different scenarios:
- Just the choices: They gave the AI only the five possible answers (A, B, C, D, E) and no question or picture. Surprisingly, the models still guessed correctly more often than random chance (which would be 20%).
- Text only: They gave the question and choices but hid the image.
- Image only: They gave the image and choices but hid the question text.
The results showed that the models prioritize text over images when processing information. When they had to choose between reading the text or looking at the image, they leaned heavily on the text. In fact, the models performed better when the image was missing (but the text was there) than when the text was missing (but the image was there). This suggests that for these specific medical questions, the AI is relying more on the words than the visual evidence.
They also broke the questions down by how much the image mattered. Some questions had pictures that were just for decoration (like a chart that repeated what the text said), while others were "image-dominant," meaning you absolutely had to look at the picture to solve it. The models struggled the most with the image-dominant questions. Even when they had the full picture and the full text, they got fewer of those right compared to the ones where the text did all the heavy lifting.
The authors suggest that this happens because the models have learned to spot patterns in the answer choices or the text descriptions that act like shortcuts, allowing them to guess correctly without truly "seeing" the medical problem. It's like a student who notices that the longest answer is usually the right one, so they just pick the longest option without reading the question.
In the end, this paper doesn't say AI is useless for medicine, but it does sound a serious alarm. It suggests that when these models get high scores on medical exams, they might not be proving they understand the visual evidence as well as we hope. They might just be really good at reading the fine print. The researchers warn that in the real world, where a doctor needs to actually see a tumor on a scan or a rash on a skin, relying on an AI that ignores the picture could be dangerous. The study concludes that while AI is getting smarter, we need to make sure it's actually looking at the evidence, not just guessing based on the clues in the text.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.