Extraction and Analysis of Multimodal Concepts in Vision Language Models through Sparse Autoencoders
This paper proposes a Sparse Autoencoder-based framework that effectively extracts and analyzes visual, textual, and multimodal concepts in Vision Language Models, significantly improving the quality of visual concept descriptions and enabling systematic identification of integrated multimodal concepts compared to existing methods.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a Vision Language Model (VLM) as a super-smart, bilingual librarian who can look at a picture and read a book simultaneously. This librarian is incredibly good at answering questions like, "What is the dog doing in this photo?" or "Write a story about this scene." However, inside this librarian's brain, there are billions of tiny switches (neurons) firing away, and until now, we had no idea what specific thoughts or ideas those switches were actually controlling. It was like watching a lightbulb flash without knowing if it meant "cat," "red," or "happy."
This paper introduces a new way to peek inside the librarian's brain using a tool called a Sparse Autoencoder (SAE). Think of the SAE as a high-tech concept sorter. Instead of looking at the messy, tangled web of the whole brain, the SAE organizes the activity into neat, separate drawers. Each drawer represents a single, clear idea (a "concept") that the librarian is thinking about.
The Problem: Missing the "Multimodal" Mix
Previous attempts to open these drawers had a major blind spot. They mostly looked at text (words) or images (pictures) separately.
- Imagine trying to understand a recipe by only reading the list of ingredients (text) or only looking at the finished cake (image), but never seeing how they work together.
- The authors argue that VLMs often have "multimodal" concepts—ideas that exist only when you combine the picture and the words. For example, a neuron might fire specifically for "a dog chasing a ball." If you only look at the picture, you see a dog and a ball. If you only read the text, you see "dog" and "ball." But the specific action of chasing is a mix of both. Previous tools missed these mixed-up ideas, leading to vague or wrong descriptions.
The Solution: A Better Detective Framework
The authors built a new framework to fix this. Here is how it works, using a simple analogy:
- The Trigger: They take a huge pile of "Question-and-Answer" cards (each with a photo and a question). They feed these to the librarian and watch which specific drawers (neurons) light up the brightest.
- The Detective (The Explainer): For each bright neuron, they pick the top 5 photos and 5 sentences that made it light up. They then ask a second, even smarter AI (an "explainer") to look at these clues and guess: "What is this neuron thinking about?"
- The Twist: Unlike old methods that just showed a blurry, masked-out picture, this new method gives the detective the full context. If the neuron is triggered by a picture, the detective also sees the question and answer. If it's triggered by text, the detective sees the image too. This helps the detective understand the relationship between the two.
- The Verdict: Once the detective guesses a concept (e.g., "a skateboarder"), the system uses a "truth checker" (models called CLIP and ALIGN) to see if that guess actually matches the data. It asks: "Does the phrase 'skateboarder' really describe these specific images?"
The Results: Sharper Focus
The paper tested this new method on a dataset of visual questions (LLaVA-NeXT) and compared it to the old way of doing things.
- Visual Quality: The new method was 45% better at correctly identifying what a visual concept was. It stopped guessing vague things and started giving precise descriptions. It was like upgrading from a blurry, low-resolution photo to a crisp, high-definition image.
- Finding the Mix: For the first time, they systematically found and labeled these "multimodal" concepts—the ones that live in the sweet spot between image and text.
- Testing the Brain: To prove these concepts actually matter, they did a "control experiment." They forced a specific neuron to stay "on" (like holding a light switch down) and asked the librarian to tell a story.
- When they forced a visual neuron (e.g., "water"), the story suddenly included water.
- When they forced a multimodal neuron, the story changed even more dramatically, showing that these mixed concepts have a powerful influence on how the model thinks and speaks.
The Bottom Line
This paper doesn't just say "we can see inside the model"; it says, "We can now see the right things inside the model." By treating images and text as partners rather than separate entities, they created a map of the librarian's brain that is much more accurate. They found that many of the most interesting ideas in these models are a blend of sight and sound, and their new tool is the first to reliably find and name them.
In short: They built a better microscope that lets us see not just the words or the pictures, but the unique ideas that only exist when the two meet.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.