← Latest papers
🤖 AI

Explanation Multiplicity: Circuit-Level Interpretability Evidence Does Not Survive Defensible Analytic Variation

This pre-registered study demonstrates that mechanistic interpretability evidence for AI regulation fails to meet filability standards because circuit discovery results are highly unstable and structurally disjoint across defensible analytic variations, rendering them unreliable for compliance documentation.

Original authors: Ajay Pravin Mahale (Hochschule Trier)

Published 2026-08-17
📖 4 min read☕ Coffee break read

Original authors: Ajay Pravin Mahale (Hochschule Trier)

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to explain a magic trick to a skeptical friend. You say, "The magician didn't just wave a wand; he used a specific hidden lever under the stage." But then your friend asks, "How do we know that lever? What if I look under a different part of the stage, or use a different flashlight, or ask a different magician to check?" If your explanation changes completely depending on how you look for the secret, can you really call it a fact? This is the heart of a new study in the world of Artificial Intelligence (AI).

In the field of "mechanistic interpretability," scientists try to peek inside AI brains (like the famous GPT-2 model) to find the specific "circuits"—tiny networks of connections—that make the AI do what it does. It's like trying to map the exact electrical wiring in a house to explain why a light turned on. The European Union has a new law saying that if an AI makes a high-stakes decision (like denying a loan), the company must file a report explaining how the AI decided. The obvious way to do this is to file a map of the AI's internal circuits. But this paper asks a tough question: If two smart experts use the same AI and the same tools, but make slightly different, reasonable choices about how to look for the circuits, will they find the same map? Or will they find two totally different stories?

The author of this paper decided to test this by setting up a massive "choose-your-own-adventure" experiment. They didn't just look at the AI once; they looked at it 15,840 different times. They created a grid of seven different ways to analyze the AI, taking every single setting from existing, published tools. They asked: "If we change the metric, the seed, the size of the circuit, or the way we test it, does the story we tell about the AI's brain stay the same?"

The answer, unfortunately, is a loud "No."

When the researchers ran their experiment, they found that the evidence is incredibly shaky. Out of all the different ways they could have analyzed the data, the story they told about the AI's circuit flipped to a completely different version 73.2% of the time. That means if two auditors looked at the same AI with the same tools but slightly different settings, they would likely file two completely different reports. Even when they tried to be super strict and standardize the most important choices, the flip rate only dropped to 59.4%.

To make sure this wasn't just a fluke, they checked if the circuits they found were actually the same thing described in different words. They weren't. The circuits were structurally almost entirely different (sharing only about 4% of their parts) and functionally uncorrelated. It's as if one auditor found a map of the kitchen wiring, and the other found a map of the plumbing, and both claimed, "This is the secret to how the house works!"

The paper also discovered that one of the seven main tools they used to find these circuits didn't work at all on the task they were testing; it crashed immediately. Furthermore, more than half of the potential explanations were thrown away because they didn't meet the strict rules of the test.

So, what does this mean for the future? The author isn't saying AI is broken or that we can't understand it. They are saying that the current "receipts" we have to prove how AI works aren't ready to be filed with a government regulator. If you hand a judge a circuit map today, and a second judge uses the same tools but a slightly different method, they might throw out your map and say, "This isn't the right explanation." The evidence, as it stands right now, fails the test of being reliable enough to be a legal document. The study suggests that before we can trust these AI explanations in court or in safety checks, we need to figure out how to make the maps stable, no matter who is holding the flashlight.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →