A Comparative Explainability Framework for DeBERTa-v3 in Zero-Shot Medical Abstract Classification
This paper presents a comparative explainability framework that audits DeBERTa-v3's zero-shot classification of medical abstracts by evaluating the agreement of five attribution methods, revealing that explanatory stability correlates with predictive certainty and identifying systemic failure modes like lexical hypersensitivity to guide future model improvements.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the modern world of artificial intelligence, computers have become remarkably skilled at reading and understanding human language. They can scan thousands of medical reports in seconds, spotting patterns that might take a human doctor hours to find. These systems, often built on complex structures called transformers, act like powerful engines that process text by looking at how words relate to one another. However, a significant problem remains: while these engines are accurate, they are often opaque. They function as black boxes, delivering a diagnosis or a classification without explaining which specific words led them to that conclusion. In the high-stakes field of medicine, where a wrong guess could affect a patient's safety, a doctor cannot simply trust a machine's output without understanding its reasoning. If a computer says a patient has a specific disease, the doctor needs to know if the machine noticed the right symptoms or if it was tricked by a single, misleading word. This need for transparency has given rise to a field dedicated to opening the black box, trying to make the internal logic of these machines visible and understandable to humans.
A team of researchers set out to test how well we can currently see inside one of these advanced language models when it is asked to classify medical abstracts without any prior training on specific medical data. They focused on a model called DeBERTa-v3, which is known for its ability to understand language nuances. The researchers did not teach the model to recognize diseases; instead, they asked it to guess the category of a medical text based on general knowledge it had already learned, a method known as zero-shot learning. To do this, they treated the task like a logic puzzle: for each medical abstract, the model was asked to decide if the text supported a specific description of a disease category. They tested this on five different groups of diseases: cancers, digestive issues, nervous system disorders, heart and blood vessel problems, and a broad catch-all category for general or systemic conditions.
The core of their investigation was to see if different methods used to explain the model's decisions would agree with one another. Imagine asking five different experts to point out the most important words in a sentence that led to a conclusion. If the experts are reliable, they should all point to roughly the same words. The researchers used five distinct techniques to generate these explanations, ranging from methods that treat the model as a mystery to be probed from the outside, to methods that look directly at the model's internal mathematical pathways. They then compared the lists of top twenty words each method highlighted for the same text. They measured how much these lists overlapped, essentially asking, "Do these different ways of looking at the model see the same thing?"
The results revealed a clear and telling pattern. When the model was classifying specific diseases like cancer or heart conditions, the different explanation methods largely agreed. They all pointed to the same key medical terms, such as "metastatic" for cancer or "liver" for digestive issues. In these cases, the model was confident, and the explanations were stable, suggesting that the system was indeed using the correct clinical logic. However, the situation changed dramatically when the model faced the broad category of general medical conditions. Here, the model's accuracy dropped significantly, and the explanation methods stopped agreeing. Instead of pointing to a shared set of medical terms, the different methods highlighted completely different words, often drifting toward common grammatical particles or unrelated terms. The lack of agreement among the explanation tools served as a warning signal that the model was struggling and that its decision was not based on a solid clinical foundation.
Through a detailed look at the errors the model made in this general category, the researchers identified three specific ways the system failed. First, the model showed a hypersensitivity to specific words. If a text contained a single strong word like "biopsy" or "malignancy," the model would immediately jump to a cancer diagnosis, even if the rest of the text described a routine procedure with no cancer context. Second, the model struggled with semantic overlap. When a text described a systemic issue that involved blood vessels, the model often confused it with a specific heart condition, unable to weigh the fact that the problem was actually affecting the whole body rather than just the heart. Finally, when the text lacked a clear, dominant medical clue, the model's explanation fell apart completely. The importance of the words scattered randomly across the sentence, with the model seemingly latching onto random grammatical words like "it" or "diagnose" rather than forming a coherent clinical reason.
The study concludes that relying on a single method to explain an AI's decision in medicine is insufficient. Because different methods can produce such different results, especially when the model is unsure, doctors and regulators need to look at the agreement between multiple explanation tools. If the tools agree, the decision is likely trustworthy. If they disagree, it is a sign that the model is confused or that the category it is trying to classify is too vague. The researchers suggest that for medical AI to be safe and auditable, we should focus on specific, well-defined disease categories rather than broad, general labels. By narrowing the scope to clear clinical definitions, we can ensure that the AI's reasoning remains stable and that the explanations it provides are actually useful for human experts.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.