← Latest papers
🤖 machine learning

Reliability, Faithfulness, and the Limits of Post-hoc Explanations of Opaque Scientific Models

The paper argues that while reliability and faithfulness are necessary conditions for interpreting scientific machine learning models, they are insufficient on their own to validate claims about the true structural mechanisms of the underlying phenomena without external corroboration.

Original authors: Nick Oh, Helen Jin

Published 2026-06-30
📖 5 min read🧠 Deep dive

Original authors: Nick Oh, Helen Jin

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Picture: The "Black Box" Detective

Imagine scientists are trying to understand a complex natural phenomenon (like how a disease works or how a star burns). They build a super-smart computer model (an AI) to predict what happens. The model is incredibly accurate—it gets the right answer almost every time. This is Reliability.

But the model is a "black box." We don't know how it got those answers. So, scientists use special tools (like SHAP or LIME) to peek inside and ask, "What clues did you use to make that decision?" The tool gives an answer, like "The model looked at the patient's age and the color of their skin." This is Faithfulness.

The Paper's Main Argument:
The authors argue that even if the model is perfectly reliable (always right) and the explanation is perfectly faithful (tells the truth about what the model did), you still cannot conclude that the model actually understands how the real world works.

You can know what the machine did, but you don't know if it did it for the right reasons.


The Three-Step Chain

The paper describes a chain of logic that scientists often try to use, which looks like this:

  1. The Real World (The Phenomenon): The actual thing we are studying (e.g., a real disease).
  2. The Model: The AI that predicts the disease.
  3. The Explanation: The tool that tells us what the AI was thinking.

The Mistake: Scientists often assume:

  • "The Model matches the Real World (Reliability)."
  • "The Explanation matches the Model (Faithfulness)."
  • Therefore: "The Explanation matches the Real World."

The Paper says: This logic is broken. The chain is missing a crucial link.


Analogy 1: The Cheating Student vs. The Genius

Imagine a student taking a math test.

  • Reliability: The student gets every single answer correct. They are a perfect predictor of the right answers.
  • Faithfulness: A proctor watches the student and reports exactly what the student did. The report says, "The student solved the problems by looking at the last digit of the question number and guessing the answer based on a pattern in the test booklet."

The Trap:
If you only look at the Reliability (perfect score) and the Faithfulness (the proctor's honest report), you might think, "Wow, the student has a brilliant new method for solving math!"

But in reality, the student is just cheating using a trick that works only on this specific test. They didn't learn the math (the structure of the phenomenon); they just found a shortcut that happens to work.

The paper argues that AI models often do this. They find "shortcuts" (like looking at the background of a photo instead of the disease) that make them accurate, but those shortcuts aren't how the real world actually works.

Analogy 2: The Two Clocks

Imagine you have two clocks.

  • Clock A is a real, high-quality clock with gears that match the movement of the sun.
  • Clock B is a broken clock that happens to be stuck at the exact time the sun is at, but only because someone set it that way.

Both clocks show the correct time (Reliability).
If you take a photo of the inside of both clocks, you get a faithful description of their insides (Faithfulness).

  • Clock A's inside shows gears turning.
  • Clock B's inside shows a frozen hand.

If you only know that both clocks show the right time, and you have a faithful description of their insides, you still cannot know which one is actually tracking the sun correctly. One of them is just lucky.

The paper says that Reliability and Faithfulness are like checking if the clocks show the right time and describing their insides. Neither check tells you if the clock is actually connected to the sun (the real mechanism).


The "Perfect" Scenario Doesn't Help

You might think, "What if the model is perfect? What if it never makes a mistake?"

The paper says: Even then, it doesn't matter.

Imagine a model that predicts a disease perfectly 100% of the time.

  • Scenario 1: The model learned the real biological cause (the "right" structure).
  • Scenario 2: The model learned a weird, invisible pattern in the data (like the hospital where the patient was treated) that happens to correlate with the disease, but isn't the cause.

Both models are 100% Reliable.
Both models have 100% Faithful explanations (telling you exactly what they looked at).

But in Scenario 2, the explanation is lying to you about the cause of the disease, even though it is telling the truth about the model. The model is right about the result, but wrong about the reason.

The Conclusion: What Can We Actually Do?

The paper concludes that we cannot use these explanations to make strong claims about how the world actually works (mechanisms) just by looking at the model and its explanation.

  • What the chain CAN do: It can give us ideas or hypotheses. It can say, "Hey, the model thinks this feature is important. Maybe we should go investigate that feature in the real world."
  • What the chain CANNOT do: It cannot prove that the feature is actually the cause.

To know the truth, scientists need external help. They need to bring in other knowledge, run real-world experiments, or use existing theories to verify if the model's "shortcut" is actually a real mechanism. The model and its explanation alone are not enough to bridge the gap between "the computer is right" and "we understand the universe."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →