← Latest papers
💬 NLP

Explanation Bias is a Product: Revealing the Hidden Lexical and Position Preferences in Post-Hoc Feature Attribution

This paper introduces a model- and method-agnostic framework to systematically evaluate and reveal hidden lexical and position biases in post-hoc feature attribution methods, demonstrating a trade-off between these biases across different models and identifying that anomalous explanations are more likely to be biased.

Original authors: Jonathan Kamp, Roos Bakker, Dominique Blok

Published 2026-02-20
📖 5 min read🧠 Deep dive

Original authors: Jonathan Kamp, Roos Bakker, Dominique Blok

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a very smart but mysterious robot (an AI) that reads sentences and decides if they are "good" or "bad." You want to know why it made that decision. So, you ask the robot to highlight the most important words it used to make up its mind.

This paper is about a group of detectives who realized that the tools used to get these highlights are biased. They don't just show you the truth; they have their own hidden preferences, like a camera that always focuses on the left side of the room or a flashlight that only shines on red objects.

Here is a simple breakdown of what they found, using some everyday analogies:

1. The Problem: The "Flashlight" Effect

Think of Feature Attribution (the method used to highlight words) as a flashlight. You shine it on a sentence to see which words the AI thought were important.

  • The Issue: If you use Flashlight A, it might highlight the word "cat." If you use Flashlight B, it might highlight the word "sat." Both flashlights are looking at the same sentence, but they give you different answers.
  • The Danger: If you don't know your flashlight is broken, you might trust the wrong explanation. The paper calls this Explanation Bias.

2. The Two Types of "Bad Habits"

The researchers discovered that these flashlights have two specific bad habits:

  • Lexical Bias (The "Favorite Word" Habit):
    Imagine a flashlight that loves punctuation. No matter what sentence you give it, it always highlights the period (.) or the comma (,) instead of the actual words. It's like a detective who only cares about the punctuation marks in a diary and ignores the story.

    • Example: One method in the study loved highlighting periods over commas, even when the period didn't matter.
  • Position Bias (The "Location" Habit):
    Imagine a flashlight that only shines on the very first or very last page of a book, ignoring the middle.

    • Example: One method (Vanilla Gradient) was obsessed with the start of the sentence. Another method (Integrated Gradient) was obsessed with the end of the sentence. They weren't looking at the content; they were just looking at where the words were.

3. The Experiment: The "Fake News" Test

To catch these bad habits, the researchers didn't use real news or complex stories. They created fake, random sentences made of nonsense.

  • The Setup: They wrote sentences like "table the table . ." or just a mix of commas and periods.
  • The Trick: Since the sentences were random, no word or position should matter. The AI shouldn't be able to guess the answer.
  • The Reveal: If the AI's "flashlight" still highlighted the first word or the period every single time, it proved the flashlight was broken. It was finding patterns that didn't exist.

4. The Findings: Old vs. New, and the Trade-Off

The researchers tested two different AI models (an older one called BERT and a newer one called ModernBERT) with six different "flashlights."

  • Newer isn't always better: You might think the newer, fancier AI model would be less biased. But the study found that the new model still had its own biases. It just had different ones.
  • The See-Saw Effect: This is the most interesting part. They found a trade-off.
    • If a method was really bad at highlighting the right words (Lexical Bias), it was usually really good at ignoring the position (Position Bias).
    • If a method was obsessed with where the words were (Position Bias), it usually ignored the specific words themselves.
    • Analogy: It's like a camera that either zooms in too much on the subject (ignoring the background) or focuses too much on the background (ignoring the subject). It's hard to get both perfect at the same time.

5. The "Outlier" Rule

The paper found a golden rule for spotting a liar: If a flashlight gives an explanation that is totally different from all the other flashlights, it's probably the broken one.

  • If 5 out of 6 flashlights point to the word "dog," and one points to the word "the," the one pointing to "the" is likely biased and untrustworthy.

6. The Takeaway for You

If you are using AI to make important decisions (like in medicine or law), and you look at the "explanation" of why the AI made a choice:

  1. Don't trust just one tool. Try different methods to see if they agree.
  2. Watch out for the "First/Last" trick. If the AI always blames the first or last word, it might be a glitch, not a real reason.
  3. Check the "Odd One Out." If one explanation looks weird compared to the others, be skeptical.

In short: AI explanations aren't perfect mirrors of reality; they are more like funhouse mirrors that stretch and shrink things based on the tool you use. This paper helps us figure out which mirrors are distorted so we don't get fooled.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →