OCCAM: Open-set Causal Concept explAnation and Ontology induction for black-box vision Models
OCCAM is a novel framework that interprets black-box vision models by discovering open-set visual concepts, quantifying their causal impact through object-level interventions, and aggregating this evidence to induce a structured ontology that reveals global model reasoning and biases.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a super-smart, but completely secretive, robot that looks at photos and tells you what's in them. You can show it a picture of a dog, and it says "Dog." But if you ask, "Why did you say that?" the robot refuses to answer. It won't show you its internal gears, its code, or its thought process. It's a "black box."
This is the problem the paper OCCAM tries to solve. The authors created a new way to peek inside this black box without breaking it open. They do this by acting like a very careful editor who cuts things out of a photo to see what happens to the robot's answer.
Here is how OCCAM works, explained through simple analogies:
1. The "What If?" Game (Open-Set Discovery)
Most previous methods tried to guess what the robot was thinking by using a fixed list of words (like a dictionary). If the robot saw a "golden retriever," but the dictionary only had "dog," it might miss the nuance.
OCCAM is different. It uses a smart AI assistant (a Multimodal Large Language Model) to look at the photo and say, "Hey, I see a fluffy tail, a sunny park, and a red ball." It doesn't need a pre-written list; it just invents the concepts it sees on the spot. This is the "Open-Set" part—it can talk about anything it finds.
2. The "Magic Eraser" (Causal Intervention)
Once OCCAM identifies these concepts (like the "fluffy tail"), it doesn't just point at them. It performs a causal test.
Imagine you have a photo of a dog in a park. OCCAM uses a "Magic Eraser" to carefully remove just the tail, filling in the background so it looks natural. Then, it shows this edited photo to the robot again.
- If the robot still says "Dog": The tail wasn't very important to the decision.
- If the robot suddenly says "I don't know" or changes its mind: The tail was a crucial clue.
By doing this for every concept it finds, OCCAM learns exactly which parts of the image are actually causing the robot to make its decision. It's like a detective removing clues one by one to see which one was the smoking gun.
3. Building a "Family Tree" of Ideas (Ontology Induction)
So far, OCCAM has explained one photo. But the authors wanted to understand the robot's entire personality, not just its reaction to one picture.
They take the results from thousands of photos and build a giant, structured map called an Ontology. Think of this as a family tree or a complex subway map for the robot's brain.
- It connects concepts to classes (e.g., "Fluffy Tail" "Dog").
- It connects concepts to other concepts (e.g., "Smiling Face" often appears with "Happy Dog").
This map isn't drawn by a human; it emerges from the robot's own behavior. It reveals hidden patterns, like "This robot always needs to see a 'smile' to classify a 'volleyball player' as happy," or "This robot is biased to think 'grass' means 'outdoor' even if the grass is fake."
4. Why This Matters
The paper shows that this method works better than previous tricks in two main ways:
- It works on secret robots: You don't need to see the robot's code (which you usually can't). You just need to be able to show it pictures and get answers.
- It finds the real reasons: By actually removing the visual evidence, it proves causality. It doesn't just guess; it tests.
The authors tested this on famous image datasets (like Broden and ImageNet) and found that OCCAM could explain the robot's decisions more accurately than methods that rely on internal code access. They also showed that when they fed this structured "Family Tree" map to another AI to summarize the robot's behavior, the summaries were much deeper and more accurate than just listing random facts.
In short: OCCAM is a tool that treats a black-box AI like a mystery to be solved by systematically erasing parts of the picture to see what the AI actually cares about, then organizing those findings into a clear map of how the AI thinks.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.