← Latest papers
💬 NLP

Legal Experts Disagree With Rationale Extraction Techniques for Explaining ECtHR Case Outcome Classification

This paper introduces a new ECtHR dataset and a comparative framework to evaluate rationale extraction techniques for legal outcome prediction, revealing that while current methods achieve high faithfulness metrics, legal experts fundamentally disagree with the extracted rationales as valid explanations for court decisions.

Original authors: Mahammad Namazov, Tomáš Koref, Ivan Habernal

Published 2026-04-08
📖 4 min read☕ Coffee break read

Original authors: Mahammad Namazov, Tomáš Koref, Ivan Habernal

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a super-smart robot lawyer. You feed it thousands of pages of court rulings, and it gets really good at guessing: "Did the human rights court find a violation in this case? Yes or No?"

The robot is accurate. But here's the problem: It won't tell you why. It just gives you a "Yes" or "No" without a clear explanation. In the real world of law, you can't just trust a guess; you need to know the reasoning.

So, researchers tried to build a "translator" for this robot. They created tools that scan the robot's brain and pull out the specific sentences it used to make its decision. These sentences are called rationales. Think of them as the robot highlighting the most important parts of a story to prove its point.

This paper is essentially a quality control check to see if these "highlighting tools" are actually any good.

The Experiment: The Robot vs. The Human Judge

The researchers set up a test with two main groups:

  1. The Robot's "Reasons": They used two different AI techniques to extract the highlighted sentences (the rationales) from the robot's decision.
  2. The Human Experts: They hired real human legal experts (lawyers who know the law inside and out) to look at those same highlighted sentences and decide: "Does this actually explain why the court found a violation?"

They also tried a third group: The AI Judge. They asked other Large Language Models (like the ones powering chatbots) to act as judges and grade the quality of the highlighted sentences, hoping to save money on hiring human lawyers.

The Big Surprise: The Robot is "Lying" (Sort Of)

Here is where it gets interesting.

The Numbers Look Great (The "Faithfulness" Trap):
When the researchers looked at the math, the robot's highlighting tools seemed perfect. The tools claimed, "Look! If we only show you these highlighted sentences, the robot still gets the right answer!"

  • Analogy: Imagine a student taking a test. The teacher asks, "Can you solve this math problem if I only show you the numbers 2 and 2?" The student says, "Yes!" The teacher checks the math, and sure enough, 2+2=42+2=4. The student seems to understand the logic.

The Human Reality Check:
But then, the human lawyers looked at the highlighted sentences. They were horrified.

  • The Problem: The robot wasn't highlighting the legal reasoning. It was highlighting random fragments.
  • The Metaphor: Imagine a detective trying to solve a murder. The human detective highlights: "The butler, the candlestick, the library, 9 PM." This makes sense.
    The robot, however, highlights: "...the butler . the . candlestick . 9 . on . the . library . 24."
    It's like the robot is spitting out a broken jigsaw puzzle. It picked up the right words (butler, library) but missed the verbs and connections that make a sentence meaningful. It's a collection of "broken snippets."

The Verdict:
The human experts said, "No, this doesn't explain anything." Even though the robot could still guess the right answer using those broken snippets, the snippets themselves were nonsense to a human. The robot's "reasons" were totally different from how a human lawyer thinks.

The "AI Judge" Attempt

The researchers also asked, "Can we just use another AI to grade these explanations instead of hiring expensive human lawyers?"

  • The Result: The AI Judges were okay at spotting the obvious nonsense (like the broken snippets), but they weren't reliable. They sometimes disagreed with each other and sometimes disagreed with the human experts.
  • The Lesson: You can't fully replace a human lawyer with a chatbot yet when it comes to judging the quality of legal reasoning. The AI is too prone to "hallucinations" (making things up) or missing the subtle context that a human catches.

Why This Matters

This paper is a wake-up call for the legal tech world.

  1. Don't trust the "Black Box" yet: Just because an AI model is good at predicting outcomes doesn't mean it understands the law. It might be finding patterns we don't see, but its "explanations" are often just garbage.
  2. Metrics can be misleading: You can have a math score that says a tool is "99% accurate," but if a human can't read the explanation, the tool is useless in a courtroom.
  3. Human oversight is non-negotiable: In high-stakes fields like law, you cannot automate the "why." You need human experts to verify that the machine's logic actually holds water.

In short: The robot is a great guesser, but a terrible explainer. Until we can fix the "explanation" part, we can't fully trust these robots to make legal decisions for us.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →