Capability Interpretability: Human Interpretability of Vision Foundation Models
This paper introduces a psychophysics-based framework demonstrating that vision foundation models are consistently less human-interpretable than supervised counterparts, revealing that interpretability is an independent dimension driven by feature locality and coarse-grained semantic alignment rather than overall model capability.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Question: Can We Understand the "Brain" of AI?
Imagine you have a super-smart robot that can identify a cat, a car, or a traffic light with incredible accuracy. It's so good at its job that it's being used to drive cars and help doctors diagnose diseases. But here's the problem: How does it actually "see" the world?
If you ask the robot, "What part of this picture made you think it was a cat?" it doesn't have a voice. It just has internal signals (features) that light up. The big question this paper asks is: Can a regular human look at those signals and understand what they mean?
The authors call this Interpretability. It's not about how smart the robot is (its capability); it's about how clear its thinking is to us.
The Old Way vs. The New Way
The Old Way (The "Multiple Choice" Trap):
Previous tests tried to see if humans understood AI features by showing them pictures and asking, "Which of these two pictures matches the AI's signal?"
- The Flaw: The authors found a trick in this method. Humans could often get the right answer just by knowing what the feature wasn't.
- Analogy: Imagine I show you a picture of a cornfield and a picture of a beach, and I say, "This AI signal is NOT about corn." You can easily pick the beach, even if you have no idea what the signal actually is. You're just eliminating the wrong answer, not truly understanding the signal. This made the test unfair and unreliable.
The New Way (The "Detective" Framework):
The authors built a new, fairer test with two parts, like a detective trying to solve a mystery:
Localizability (The "Where" Game):
- The Setup: You are shown a "mystery signal" along with a few examples of what triggers it (e.g., a picture of a tire, a heatmap showing where the signal is bright).
- The Task: You are shown a new picture and asked to click exactly where you think that signal would light up.
- The Goal: Can you point to the specific part of the image (like the tire) that matches the signal?
Nameability (The "What" Game):
- The Setup: Same visual clues as above.
- The Task: You have to write a short description of what the signal represents.
- The Goal: Can you describe it in words? (e.g., "It's a car wheel" or "It's a red traffic light").
To make sure they were testing pure "features" and not just messy, mixed-up neurons, they used a special tool (a Sparse Autoencoder) to isolate single, clear concepts, like separating individual ingredients from a smoothie.
The Shocking Discovery: Smarter Doesn't Mean Clearer
The researchers tested six different AI models:
- 2 Older Models: Trained in a traditional, supervised way (like a student with a teacher).
- 4 "Foundation" Models: The new, massive, super-powerful models (like DINO, CLIP) that are trained on the whole internet and can do almost anything.
The Result:
The "Foundation" models were less interpretable than the older, simpler models.
- The Analogy: Imagine the older models are like a student who learned by memorizing a textbook. You can easily ask them, "Why did you pick that answer?" and they give a clear reason.
- The new Foundation models are like a genius who learned by reading the entire library of human knowledge. They are smarter and can solve harder problems, but if you ask them to explain their reasoning, they give you a vague, confusing answer. Their internal "thoughts" are harder for humans to pin down.
Crucially, being "smarter" didn't help. The paper found zero connection between how well the model performed on tasks (like identifying objects) and how easy it was for humans to understand its features. Being a better driver doesn't mean your brain is easier to read.
What Makes a Feature Easy to Understand?
The authors found two specific things that make an AI's "thoughts" easier for humans to grasp:
Locality (The "Spotlight" Effect):
- If a feature lights up in a tight, specific spot (like just the tire of a car), humans can understand it easily.
- If a feature lights up everywhere at once (like the whole car and the road and the sky), humans get confused. The new Foundation models tend to have these "diffuse" signals that cover too much ground.
- Analogy: It's easier to find a specific person in a crowd if they are wearing a bright red hat (local) than if they are just "part of the general vibe" of the crowd (diffuse).
Coarse-Grained Alignment (The "Big Picture" Match):
- Humans understand the world in broad categories (e.g., "Animals," "Vehicles"). If the AI organizes its features to match these big categories, humans understand it better.
- If the AI focuses on tiny, hyper-specific details (like the exact texture of a butterfly's wing compared to another butterfly), humans actually find it harder to understand.
- Analogy: It's easier to explain a map if you show me "North America" and "South America" (coarse) rather than trying to explain the specific shape of every single street in a city (fine-grained).
The One Bright Spot: DINOv3
There was one exception. A model called DINOv3 was both very capable and very interpretable. Why? Because the people who built it specifically programmed it to look for "local" features (things that light up in specific spots). This proves that we can build smart models that are also easy to understand, but we have to design them with that goal in mind.
Summary
- Capability Interpretability: Just because an AI is super smart doesn't mean we can understand how it thinks. In fact, the newest, smartest models are often the hardest to understand.
- The Gap: There is a gap between "doing the job well" and "explaining the job clearly."
- The Fix: To make AI more transparent, we need to train them to focus on specific, local details and organize their knowledge in broad, human-like categories, rather than just making them bigger and more powerful.
The paper concludes that we cannot assume that as AI gets smarter, it will naturally become more understandable. We have to build that clarity in from the start.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.