← Latest papers
💬 NLP

Mechanistic Interpretability as Statistical Estimation: A Variance Analysis

This paper argues that mechanistic interpretability is fundamentally a statistical estimation problem plagued by high intrinsic variance in causal mediation scores, which causes circuit discovery pipelines to produce unstable and fragile results that necessitate a shift toward rigorous statistical robustness and stability reporting.

Original authors: Maxime Méloux, François Portet, Maxime Peyrard

Published 2026-06-01
📖 5 min read🧠 Deep dive

Original authors: Maxime Méloux, François Portet, Maxime Peyrard

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Picture: Trying to Map a City with a Shaky Compass

Imagine you are trying to draw a map of a complex city (a neural network) to understand how it gets from Point A to Point B (how it solves a problem). Scientists in the field of Mechanistic Interpretability are like urban explorers. They want to find the specific "roads" and "intersections" (neurons and connections) that the city uses to make decisions. They call these special routes "circuits."

The paper argues that right now, these explorers are using a compass that spins wildly. They think they are finding the one true map, but in reality, they are drawing different maps every time they look, depending on tiny, random changes in the weather or how they hold the compass.

The Core Problem: The "Spinning Compass" (Variance)

The researchers discovered that the tools used to find these circuits are fundamentally unstable. They broke the problem down into three layers of instability:

1. The Raw Signal is Noisy (Intrinsic Variability)

The Analogy: Imagine trying to measure the strength of a specific wind current in a storm. If you measure it at one exact second, you get a number. If you measure it one second later, the wind might be slightly stronger or weaker because of random turbulence.
The Paper's Claim: The "importance" of a specific part of the AI isn't a fixed number like the weight of a rock. It's more like the wind speed. It changes wildly depending on the specific input (the "storm") you are looking at. The paper shows that even if you calculate this perfectly, the score for a single connection fluctuates so much that it's hard to tell if it's actually important or just a random gust.

2. The Tools Add More Noise (Approximation Errors)

The Analogy: Because measuring the wind perfectly takes too much time, the explorers use a cheap, fast sensor (approximation methods like EAP) to guess the wind speed. But this cheap sensor is jittery. It adds its own static and errors on top of the already changing wind.
The Paper's Claim: The fast methods scientists use to save time (like Edge Attribution Patching) introduce even more noise. They often make the scores look more unstable than they already were. The "signal-to-noise ratio" is so low that the tools are often just picking up static rather than the real signal.

3. The Map Changes with Every Step (Aggregation Instability)

The Analogy: To draw the final map, the explorers take 100 different measurements and average them out. But because the measurements are so shaky, if you change the list of 100 measurements just a little bit (like swapping one for another), the final map looks completely different. One day, the map shows a highway; the next day, it shows a dirt path.
The Paper's Claim: When scientists combine these shaky scores to build a "circuit," the result is incredibly fragile.

  • Data Changes: If they use a slightly different set of examples to train their map, they find a totally different circuit.
  • Method Changes: If they tweak a setting (like changing how they average the numbers), they find a different circuit.
  • Perturbation Changes: If they change how they "break" the AI to test it (the counterfactual), the circuit they find shifts again.

The "Dead Salmon" Effect

The paper mentions that current methods are prone to "dead salmon" artifacts.
The Analogy: In neuroscience, researchers once found "brain activity" in a dead salmon just because they didn't correct for statistical noise. The paper suggests AI researchers might be doing the same thing: finding "circuits" that look real but are actually just random noise that happened to line up by chance.

What the Researchers Found (The Evidence)

They ran experiments on different AI models (like GPT-2 and Llama) and tasks (like math or grammar).

  • The Result: When they ran the same experiment 100 times with slightly different data, the "circuits" they found barely overlapped. Sometimes, the overlap was less than 10%.
  • The Conclusion: There is no single "True Circuit" waiting to be discovered. Instead, there is a cloud of possible circuits, and current methods are just picking one random point from that cloud and calling it the truth.

The Proposed Solution: Stop Guessing, Start Measuring Uncertainty

The authors don't say "give up on AI interpretation." Instead, they say "stop pretending we know more than we do."

The New Rulebook:

  1. Report the Variance: Don't just show one map. Show 100 maps and say, "Here is the average, and here is how much they differ."
  2. Check Stability: Before claiming you found a circuit, check if it stays the same when you change the data slightly. If the circuit disappears with a tiny change, it's not reliable.
  3. Embrace Uncertainty: Treat these findings like weather forecasts (e.g., "70% chance of rain") rather than absolute facts (e.g., "It will rain").

Summary

The paper argues that Mechanistic Interpretability is currently trying to do science without admitting it's dealing with statistics. It's like trying to measure the height of a mountain while standing on a trampoline. The authors want the field to stop looking for a single, perfect "ground truth" circuit and start reporting how shaky and variable their findings actually are. Only by acknowledging this instability can we trust the maps we are drawing of AI minds.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →