← Latest papers
💬 NLP

Position: Mechanistic Interpretability Must Disclose Identification Assumptions for Causal Claims

This paper argues that mechanistic interpretability research frequently conflates validation metrics with causal identification by failing to explicitly state the necessary assumptions, and it proposes a new disclosure norm requiring authors to clearly articulate their identification strategies and the robustness of their causal claims.

Original authors: Zezheng Lin, Fengming Liu

Published 2026-05-11
📖 6 min read🧠 Deep dive

Original authors: Zezheng Lin, Fengming Liu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to figure out why a complex machine (an AI) does what it does. You suspect a specific gear (a "circuit" or "feature") is the cause of a specific action. You pull that gear out, and the machine stops working. You declare, "Aha! That gear was the cause!"

This paper argues that in the field of AI "mechanistic interpretability," researchers are making these declarations far too often without following the rules of evidence. They are claiming to have found the cause, but they haven't explained why their experiment proves it's the cause and not just a coincidence.

Here is the breakdown of the paper's argument using simple analogies:

1. The Core Problem: "Validation" is not "Identification"

The paper makes a crucial distinction between two things:

  • Validation (The "What"): "I pulled the gear, and the machine stopped." This is a fact. It's a successful test.
  • Identification (The "Why"): "I know for a fact that only that gear caused the machine to stop, and not some hidden backup gear or a side effect of my pulling."

The Analogy:
Imagine you are a doctor. You give a patient a pill, and their fever goes down.

  • Validation: "The fever went down after the pill." (This is true).
  • Identification: "The pill caused the fever to go down." (This requires assumptions).

Maybe the fever went down because it was the time of day the fever naturally breaks. Maybe the patient drank water at the same time. Maybe the pill didn't work, but the fever was going away anyway.

In AI research, papers often say, "We changed this part of the AI, and the behavior changed, so this part is the cause." The paper argues they are skipping the step of proving that no other explanation could have caused that change. They are treating the "fever going down" (validation) as proof of the "pill working" (identification) without checking if the patient had a natural fever cycle.

2. The "Hidden Assumptions" Trap

Every time a researcher claims to have found a cause, they are standing on a set of invisible assumptions. The paper says these assumptions are currently hidden.

The Analogy:
Imagine a magician claims, "I made the rabbit disappear because I waved my wand."

  • The Hidden Assumption: "The rabbit didn't jump out the back door."
  • The Hidden Assumption: "The rabbit wasn't already gone before I waved."

In AI papers, the "magic tricks" are things like:

  • Activation Patching: "We swapped this part of the AI's brain. The behavior changed."
    • Hidden Assumption: "We didn't accidentally wake up a sleeping backup system that also does this job."
  • Sparse Autoencoders (SAEs): "We found a specific 'feature' in the AI that represents the concept of 'honesty'."
    • Hidden Assumption: "This feature is a unique, standalone building block, not just a messy mix of other things that happens to look like 'honesty'."

The paper audits 10 major papers (and checks 30 more) and finds that zero of them have a dedicated section where they list these hidden assumptions. They just show the magic trick (the result) and say "Look, it works!" without admitting what rules they are assuming the universe follows.

3. The "Substitution" Error

The paper calls the current habit "Validation Metric Substitution."

The Analogy:
Imagine a car mechanic says, "I fixed your engine because the car started."

  • The Error: The mechanic is substituting the result (the car started) for the proof (I know exactly which part I fixed and that nothing else could have started the car).
  • The Reality: The car might have started because the battery recharged itself, or the mechanic bumped the key, or the engine was fine all along.

In AI papers, researchers report metrics like "faithfulness" (does the explanation match the model?) or "completeness" (did we find all the parts?). The paper argues these are just "the car started" moments. They are good to have, but they are not proof of causality unless you state the assumptions that make them proof.

4. The Proposed Solution: The "Disclosure Protocol"

The author suggests that the field should borrow a rule from Econometrics (the study of economics data). In economics, if you claim "X causes Y," you must explicitly state:

  1. What you are claiming: "We claim X causes Y."
  2. The Strategy: "We are using method Z to prove it."
  3. The Assumptions: "We assume that [Condition A] and [Condition B] are true."
  4. The Stress Test: "If [Condition A] is false, our conclusion falls apart. Here is how we tested if it's true."

The Analogy:
Instead of just saying "The pill worked," the doctor must write a report:

  • "We claim the pill lowered the fever."
  • "We assume the patient didn't take other meds."
  • "We assume the fever wasn't naturally ending."
  • "If the patient did take other meds, our conclusion is wrong. Here is the data showing they didn't."

5. What the Audit Found

The author looked at 10 key papers in the field (covering different methods like circuit discovery and autoencoders) and then checked 30 more.

  • Result: 0 out of 30 papers had a dedicated section listing their identification assumptions.
  • Result: Most papers (about 19 to 26 out of 30, depending on how strict you are) used "validation metrics" (like "the AI behaved as expected") as a substitute for proving their assumptions were true.
  • Result: Even the most careful papers, which found errors in others' work, didn't explicitly list the assumptions they were using to find those errors.

Summary

The paper is a call for honesty and clarity. It doesn't say AI research is wrong; it says AI research is incomplete.

Right now, researchers are saying, "Look, we found the cause!"
The paper says, "Please stop and tell us what rules you are assuming to make that claim. If you don't tell us the rules, we don't know if your 'cause' is real or just a lucky guess."

It's asking the field to move from "Look how cool our magic trick is" to "Here is exactly how the trick works, and here is why we know it's not a trick."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →