A Monosemantic Attribution Framework for Stable Interpretability in Clinical Neuroscience Transformer-Based Language Models
This paper introduces a unified interpretability framework that leverages monosemantic feature extraction to generate stable, input-level importance scores for Transformer-based language models, thereby addressing the challenges of polysemanticity and inter-method variability to enable trustworthy clinical applications in Alzheimer's disease diagnosis.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Problem: The "Noisy Radio"
Imagine you have a very smart radio (a Large Language Model) that can listen to a patient's medical history and tell you if they might have Alzheimer's disease. The radio works great at making the prediction, but it's terrible at explaining why.
When you ask the radio, "Why did you say this patient has Alzheimer's?" it gives you a confusing answer. It's like the radio is tuned to a station where all the voices are talking over each other at once. In the paper's language, the radio's internal parts (neurons) are polysemantic. This means one single "voice" in the radio is trying to talk about "memory loss," "age," and "blood pressure" all at the same time. Because everything is mixed up, the explanations the radio gives are unstable, confusing, and sometimes wrong. Doctors can't trust a radio that changes its story every time you ask it a question.
The Solution: The "Translator" (SAE)
The authors built a special translator called a Sparse Autoencoder (SAE). Think of this translator as a "de-mixer" for the radio.
- Before the Translator: The radio's internal voices are a chaotic jam session.
- After the Translator: The translator takes that chaotic noise and separates it into distinct, single-note instruments. Now, one voice only talks about "memory," another only talks about "walking speed," and another only talks about "family history."
The paper calls this monosemanticity. By forcing the radio to speak in these clear, single-concept notes, the explanations become much more stable and easier to understand.
The "Group Chat" Problem
Even with the translator, the authors noticed a new problem. They tried six different ways to ask the radio for an explanation (like asking six different people to summarize a movie). Sometimes Person A said the movie was about love, and Person B said it was about war. They didn't agree. In the paper, this is called inter-method variability. If the explanation changes depending on which "person" you ask, doctors can't trust it.
The "Editor" (TEO)
To fix the disagreement, the authors created an Explanation Optimizer (TEO). Think of this as a smart editor who sits in the middle of the group chat.
- The editor listens to all six different explanations.
- It figures out which parts of the story are consistent and which parts are just noise.
- It writes a single, perfect summary that combines the best parts of all six, ensuring the story is stable and makes sense.
The "Team Photo" (UMAP Constraint)
Finally, the authors wanted to make sure that when they looked at the explanations for many patients at once, the patterns looked organized.
Imagine taking a photo of a crowd. Without guidance, everyone might be standing in a messy, random pile. The authors added a linear constraint (a rule) that acts like a photographer shouting, "Everyone, line up in a straight row!"
This ensures that when they look at the data for a whole group of patients, the patterns line up neatly. This helps them spot the "big picture" trends that apply to most people, rather than getting lost in individual quirks.
What They Found
The paper tested this system on real medical data from two different groups of patients (one from the US and one from Latin America).
- The Result: When they used the "Translator" (SAE) and the "Editor" (TEO), the explanations became stable. The radio stopped changing its story.
- The Trade-off: Sometimes, making the explanation super stable made it a little less "sparse" (meaning it highlighted a few more words than strictly necessary), but the authors found a way to balance this so the explanations were both stable and focused.
- The Clinical Win: The system successfully identified specific medical tests (like memory tests or questionnaires about daily activities) that were most important for diagnosing Alzheimer's. It did this consistently, even when looking at different groups of patients.
The Bottom Line
The paper doesn't claim to cure Alzheimer's. Instead, it claims to fix the flashlight doctors use to see how AI makes decisions. By untangling the messy internal signals of the AI and organizing the explanations, they created a tool that gives doctors a clear, stable, and trustworthy view of why an AI thinks a patient might have the disease. This makes it safer to use these powerful AI tools in real hospitals.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.