← Latest papers
🤖 AI

Actionable Interpretability Must Be Defined in Terms of Symmetries

This paper argues that actionable AI interpretability must be formally defined through four specific symmetries, which enable the rigorous testing, design, and verification of interpretable models as a subclass of probabilistic systems capable of unifying inference, alignment, and safety compliance.

Original authors: Pietro Barbiero, Mateo Espinosa Zarlenga, Francesco Giannini, Alberto Termine, Filippo Bonchi, Mateja Jamnik, Giuseppe Marra

Published 2026-06-15
📖 5 min read🧠 Deep dive

Original authors: Pietro Barbiero, Mateo Espinosa Zarlenga, Francesco Giannini, Alberto Termine, Filippo Bonchi, Mateja Jamnik, Giuseppe Marra

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Problem: "What Does 'Understandable' Even Mean?"

Imagine you have a robot that predicts the weather. It's incredibly accurate, but when you ask it why it thinks it will rain, it just says, "Because of the math." You can't understand the math, so you can't trust the robot.

For years, researchers have tried to build "interpretable" AI (AI that humans can understand). But this paper argues that the whole field is broken. Why? Because everyone defines "understandable" differently. Some say it means "simple," others say "transparent," and others say "explainable." Because there is no single, strict rulebook, we can't actually test if a model is truly interpretable or just pretend to be.

The authors say: We need a new rulebook based on "Symmetries."

In physics, a "symmetry" is when you change something (like rotating a snowflake) and it still looks the same. The authors argue that for an AI to be interpretable, it must stay "the same" (or predictable) even when we change how we look at it or who is looking at it. They propose four specific symmetries that act as a checklist. If a model passes all four, it is truly interpretable.


The Four Symmetries (The Four Rules)

Think of these rules as a way to ensure the AI speaks the same language as a human and follows the same logic.

1. Inference Equivariance: "The Translation Test"

The Idea: If you translate the AI's output into human language, a human should be able to predict what the AI will say before it says it.
The Analogy: Imagine a secret code. If you give a human a "decoder ring" (a translation), they should be able to guess the message the AI is about to send. If the human can't guess the result even with the translation, the AI isn't interpretable.
The Paper's Claim: This rule forces the AI to be predictable. If a human can't simulate the AI's thinking process in their head, the AI fails this test.

2. Information Invariance: "The Trash Can Rule"

The Idea: The AI should only keep the information that matters and throw away the rest.
The Analogy: Imagine you are trying to identify a dog. You need to know it has four legs and fur. You don't need to know the exact shade of red on the owner's shirt in the background.
The Paper's Claim: A truly interpretable model acts like a smart filter. It discards the "noise" (irrelevant pixels or data) and keeps only the "signal" (the features needed to make the decision). If the model keeps every tiny detail, it's too messy to understand.

3. Concept-Closure Invariance: "The Vocabulary Match"

The Idea: The AI must use concepts that humans actually use and understand.
The Analogy: Imagine the AI says, "The object is 'Glorp'." If "Glorp" isn't a word humans know, you can't understand it. But if the AI says, "The object is 'Red' and 'Round'," and you know what "Red" and "Round" mean, you understand it.
The Paper's Claim: The AI's internal "concepts" must match human concepts perfectly. If the AI uses a weird internal label that doesn't map to a human idea, it fails. This ensures the AI isn't just using a secret code; it's using our shared vocabulary.

4. Structural Invariance: "The Mental Model Match"

The Idea: The AI's internal logic must match the way the human thinks.
The Analogy: Imagine a student who only understands simple addition. If you show them a complex calculus equation, they can't understand it, even if the equation is "correct." However, if you show them a simple addition problem, they can.
The Paper's Claim: An AI is only interpretable to you if its structure fits your brain. If you are a human who thinks in straight lines (linear logic), the AI must be a straight line. If the AI is a tangled knot of complex math, it's not interpretable to you, even if it's smart.


The Solution: A "Recipe" for Building AI

The authors don't just list problems; they offer a "recipe" (a mathematical framework called a Category) to build AI that passes these tests.

  • String Diagrams: They use visual diagrams (like circuit boards) to show how these AI models are built. Instead of a black box, you can see the wires and boxes.
  • The Magic Ingredient: By following these four symmetry rules, the AI becomes a "probabilistic model" that is mathematically guaranteed to be interpretable.

Why This Matters: The "Three Magic Tricks"

Once you have an AI built with these symmetries, you can perform three powerful "magic tricks" that are impossible with standard "black box" AI:

  1. Alignment (Teaching): You can mathematically prove the AI's concepts match human concepts. It's like checking if the AI's dictionary is the same as yours.
  2. Intervention (Tweaking): You can ask, "What if I change this specific concept?" and the AI will tell you the result. It's like turning a knob on a machine and seeing exactly what happens.
  3. Counterfactuals (What If?): You can ask, "What would have happened if the input was different?" The AI can simulate this "alternate reality" logically.

The Bottom Line

The paper argues that we stop guessing what "interpretability" means. Instead, we should build AI that satisfies these four strict symmetry rules. If an AI passes these tests, it is proven to be understandable, predictable, and safe to use. If it doesn't pass, it's just a black box, no matter how much it claims to be "explainable."

In short: Don't just ask the AI to explain itself. Build the AI so that its structure forces it to be understandable by matching human logic, vocabulary, and information needs.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →