Explaining BiomedCLIP with Weighted Banzhaf Interactions Supported by Tree-Gram Parsing
This paper introduces ParseFIxLIP, a novel explanation framework that integrates Tree-Gram Parsing with Weighted Banzhaf interactions to group fragmented text tokens into semantically coherent medical concepts, thereby enhancing the interpretability and robustness of BiomedCLIP's decision-making in clinical settings.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a super-smart robot how to understand the world by showing it pictures and reading stories to it at the same time. This is the world of Vision-Language Models (VLMs), a branch of artificial intelligence that acts like a digital detective, connecting what it sees in an image with what it reads in a caption. In the medical field, these robots are being trained to look at X-rays and CT scans and read the doctors' reports to help diagnose patients. But here's the catch: these robots are often "black boxes." They give an answer, but they don't explain why. If a robot says, "This X-ray shows a broken bone," a doctor needs to know exactly which part of the image and which specific words in the report led to that conclusion. Without a clear explanation, doctors can't trust the robot in life-or-death situations.
To solve this mystery, scientists use a mathematical tool called Game Theory. Think of the AI's decision-making process as a team game where every single piece of information (a pixel in the image or a tiny chunk of a word) is a "player." The game asks: "If we remove this player, does the team's score drop?" By playing this game over and over, we can figure out which players are the MVPs and which are just sitting on the bench. However, there's a major problem with how these games are currently played in medicine. The AI doesn't read words like "heart" or "fracture" as whole units. Instead, it chops them up into tiny, meaningless fragments (like breaking "heart" into "he" and "art"). This is like trying to solve a jigsaw puzzle where someone has cut the pieces into dust before you even started. The result is a chaotic, confusing explanation that no human can understand.
This is where a new study by Jakub Rymarski and his team comes in. They are tackling the problem of how to make these AI explanations clear and trustworthy for doctors. They introduce a clever new method called ParseFIxLIP. Instead of letting the AI play the game with those tiny, broken word fragments, they use a linguistic map (called a dependency parser) to glue the fragments back together into meaningful chunks before the game starts. Imagine taking those dust particles and gluing them back into whole words, and then gluing those words into full phrases like "saddle embolus" or "right kidney." By doing this, the game becomes much simpler and the answers become much clearer.
The team tested their idea on a medical AI called BiomedCLIP using a dataset of radiology images and reports called ROCOv2. They compared their new "glued-together" method against the old "broken-fragments" method. The results were striking. When the medical reports were short, both methods worked okay, but the new method was more stable. However, when the reports got long and complex (over 30 words), the old method completely fell apart. It tried to juggle too many tiny pieces at once, leading to a "curse of dimensionality" where the math got so messy that the explanation became useless (even showing negative scores, which means the explanation was worse than random guessing). In contrast, the new ParseFIxLIP method kept its cool. By grouping words into smart, semantic units (like treating "saddle embolus" as one single player instead of four broken ones), it reduced the complexity of the game. This allowed the AI to give a clear, statistically robust explanation even for long, complicated medical reports.
The researchers also ran some fun "stress tests" to see how the AI really thinks. They found that the AI is surprisingly fragile. If you rotate a medical image upside down, the AI gets confused and stops recognizing the anatomy, even though the shape is the same. It also tends to "cheat" by focusing too much on obvious markers like red arrows or text labels on the image, rather than the actual medical condition. ParseFIxLIP was able to reveal these quirks clearly, showing exactly where the AI was looking and what it was ignoring.
In short, this paper suggests that by using the structure of language to group words together before asking the AI to explain itself, we can get much better, more reliable answers. It doesn't just make the math work better; it makes the AI's reasoning look more like how a human doctor thinks—by understanding whole concepts rather than scattered fragments. While the method isn't a magic fix for every problem (it still relies on the underlying language tools, which can sometimes make mistakes), it offers a much clearer window into how these powerful medical AIs are making their decisions, which is a huge step toward trusting them in real hospitals.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.