← Latest papers
🤖 machine learning

Measuring the (Un)Faithfulness of Concept-Based Explanations

This paper introduces Surrogate Faithfulness (SURF), a new benchmark that exposes flaws in existing metrics for concept-based explanations and reveals that many state-of-the-art unsupervised methods are less faithful to model computations than previously claimed.

Original authors: Shubham Kumar, Narendra Ahuja

Published 2026-03-31
📖 5 min read🧠 Deep dive

Original authors: Shubham Kumar, Narendra Ahuja

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a brilliant but mysterious chef (the AI model) who can cook incredible dishes. You ask the chef, "Why did you put salt in this soup?"

The chef doesn't speak your language. Instead, they point to a complex, glowing map of their brain and say, "Because of the activation in neuron #4,592." That's not very helpful to you.

Concept-Based Explanations (CBEMs) are like hiring a translator. The translator looks at the chef's brain map and says, "Ah, the chef used salt, pepper, and fresh basil." Suddenly, the explanation makes sense to a human.

But here is the problem: Is the translator actually telling the truth? Or are they just making up a story that sounds nice, even if it doesn't match what the chef actually did?

This paper, "Measuring the (Un)Faithfulness of Concept-Based Explanations," is like a detective investigating whether these translators are lying.

The Big Problem: The "Fake" Translator

The authors found that the current best translators (called U-CBEMs) are actually quite bad at telling the truth, even though they look good. They found two main tricks these translators use to fool us:

  1. The "Over-Engineered" Backdoor: Some translators use a super-complex, secret machine to turn "basil" back into the final soup taste. They claim, "See? If we put basil in, we get the right soup!" But that secret machine is so complicated that you (the human) can't understand how it works. So, the explanation isn't actually simple; it just hides the complexity behind the scenes.
  2. The "Delete and See" Game: Other translators say, "If I take the 'basil' out of the soup, the taste changes a lot, so basil must be important!" The authors show that this game is rigged. Removing ingredients in a computer simulation often creates "fake soup" (data that doesn't exist in reality), and the chef's reaction to fake soup tells us nothing about their real cooking.

The Solution: SURF (The Honest Translator)

The authors propose a new way to test these translators called SURF (Surrogate Faithfulness).

Think of SURF as a simple, honest accountant.

  • No Magic Tricks: Instead of a complex machine, SURF uses a simple calculator (a linear equation). It says, "If the chef used 2 units of salt and 1 unit of pepper, the soup should taste exactly like this."
  • Checking the Whole Menu: Old tests only checked if the translator got the one right dish (e.g., "Did they explain why the soup was salty?"). SURF checks the entire menu. "Did they explain why the soup wasn't sweet? Why it wasn't spicy? Why it wasn't burnt?" If the translator gets the main dish right but gets the side dishes wrong, SURF catches them.

The "Sanity Check" (The Lie Detector Test)

To prove their new method works, the authors did a simple test. They took a perfect translator and started randomizing it.

  • Scenario A: The translator uses the real ingredients (Salt, Pepper).
  • Scenario B: The translator uses random numbers (Salt, "Clouds", "Tuesday").
  • Scenario C: The translator uses completely random nonsense.

The old methods failed this test. They sometimes said the "Random Nonsense" translator was more faithful than the real one! It was like a lie detector saying, "This guy is telling the truth!" when he was speaking gibberish.

SURF passed. It correctly said, "The real ingredients explain the soup perfectly. The random nonsense explains nothing."

The Shocking Discovery

When the authors applied SURF to the "State-of-the-Art" (the best) translators currently in use, they found a harsh truth: Most of them are not faithful.

They are like a magician who makes a rabbit appear from a hat. The audience (humans) sees the rabbit (the concept) and thinks, "Wow, that's how the trick works!" But the magician is actually using a hidden trapdoor (the model's internal math) that the rabbit has nothing to do with. The explanation looks great, but it doesn't reflect reality.

Why This Matters

In high-stakes fields like healthcare or finance, we can't trust a doctor or a loan officer if they are using a "magic trick" to explain their decisions.

  • If a doctor says, "I prescribed this medicine because the patient has a 'red rash' (a concept)," but the AI actually prescribed it because of a hidden pattern in the patient's age and zip code, the explanation is dangerous.
  • SURF gives us a way to catch these dangerous lies. It forces AI developers to stop making pretty, fake explanations and start building honest ones.

The Takeaway

The paper concludes that while we love explanations that are easy to understand, we shouldn't accept them if they are lies. SURF is the new ruler we need to measure if an AI's explanation is actually true, ensuring that when an AI says, "I did this because of X," it really means it.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →