← Latest papers
💬 NLP

A Systematic Comparison between Extractive Self-Explanations and Human Rationales in Text Classification

This paper systematically compares extractive self-explanations from instruction-tuned LLMs with human rationales across multiple text classification tasks, finding that while alignment between the two depends on text length and task complexity, self-explanations provide faithful token-level insights that differ fundamentally from the structural focus of post-hoc attribution methods.

Original authors: Stephanie Brandl, Oliver Eberle

Published 2026-05-21
📖 5 min read🧠 Deep dive

Original authors: Stephanie Brandl, Oliver Eberle

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a very smart, but sometimes mysterious, robot assistant. You ask it a question, and it gives you an answer. But now, the robot also says, "Here is why I chose that answer."

This paper is like a detective story where the researchers put these robot assistants on trial to see if their "reasons" are actually honest and helpful. They wanted to know: When a robot explains itself, is it telling the truth about how it thinks, or is it just making up a story that sounds nice?

Here is the breakdown of their investigation using simple analogies:

1. The Three Test Drives

The researchers didn't just ask the robots one type of question. They gave them three very different "driving tests" to see how they handled different terrains:

  • The Movie Review Test (Sentiment): Short, simple sentences like "This movie was great!" or "I hated the ending." This is like a quick drive on a smooth, straight road.
  • The Detective Report Test (Forced Labor): Long, complex news articles about serious human rights issues. This is like driving through a dense, foggy forest where you have to spot hidden clues (like "abusive working conditions") that aren't always obvious.
  • The Fact-Checker Test (Climate Claims): Checking if a statement about climate change is true or false based on a pile of Wikipedia snippets. This is like a puzzle where the pieces don't always fit perfectly, and you have to figure out which piece actually matters.

2. The Two Types of "Reasons"

The researchers compared two ways of getting an explanation:

  • The Robot's Own Story (Self-Explanation): The robot is asked, "Why did you pick this answer?" and it writes a sentence explaining itself. It's like a student raising their hand and saying, "I picked 'True' because the text mentioned 'ice melting'."
  • The Human's Highlighter (Human Rationale): Real humans read the same text and highlight the specific words that made them decide. This is the "gold standard" or the "truth" the researchers use to check the robots.
  • The "X-Ray" Machine (Post-hoc Attribution): This is a third method where researchers use math tools to look inside the robot's brain and see which words lit up the most during processing. It's like an X-ray showing which bones are under stress, regardless of what the robot says it was thinking.

3. The Big Findings

A. The "Length" Problem
The robots were great at explaining the short movie reviews. They matched human highlights almost perfectly. But as the text got longer and more complex (like the news articles), the robots started to wander off. Their explanations became less like what a human would pick and more like a guess.

  • Analogy: It's easy for a robot to explain a one-sentence joke. But if you ask it to explain a whole novel, it starts to get confused about which plot points actually mattered.

B. The "Truth" vs. The "Story"
Here is the most surprising part. The researchers tested if the robot's explanation actually caused the answer.

  • The Robot's Story: When the researchers covered up the words the robot claimed were important, the robot's answer changed drastically. This means the robot's explanation was faithful—it was actually using those words to make its decision.
  • The X-Ray Machine: When they used the mathematical "X-ray" to find important words, they found something weird. The X-ray kept pointing at boring, structural words like "The," "Start," or "End of text."
  • Analogy: Imagine a chef says, "I made this soup because of the salt and pepper." If you take away the salt and pepper, the soup tastes different. That's a faithful explanation. But the "X-ray" machine would say, "No, the most important thing is the pot the soup is in!" The robot's story is actually more honest about what it used to think, even if the X-ray says the pot is important.

C. The Style Difference

  • Humans tended to highlight words that told a story or described feelings (e.g., "vulnerable," "victims," "poor").
  • Robots tended to highlight words that sounded like facts or data (e.g., "hours," "dollars," "ILO," "statistics").
  • The X-Ray tended to highlight the "metadata" (dates, source names, URLs).

4. The Conclusion

The paper concludes that while robots are getting better at explaining themselves, their ability to do so depends heavily on how hard the task is.

  • Good News: When a robot gives you a self-explanation, it is often actually using those words to make its decision. It's not just lying; it's genuinely pointing to the right clues.
  • Bad News: If you use traditional mathematical tools (X-rays) to see what the robot is thinking, you might get misled. Those tools often point at the "formatting" of the text rather than the actual meaning.

In short: If you want to know why a robot made a decision, listening to its own story (self-explanation) is often more reliable than trying to take an X-ray of its brain, especially for short, clear tasks. But for long, complex stories, even the robots get a bit lost in their own explanations.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →