← Latest papers
🤖 AI

Where Experts Disagree, Models Fail: Detecting Implicit Legal Citations in French Court Decisions

This paper introduces a benchmark for detecting implicit citations of the French Civil Code and demonstrates that expert disagreement on these cases serves as a critical signal of intrinsic difficulty, revealing that while models struggle with disputed instances, a multi-model consensus approach can effectively rank high-confidence candidates without supervision.

Original authors: Avrile Floro (UPHF), Tamara Dhorasoo (UPHF), Soline Pellez (UPHF), Nils Holzenberger

Published 2026-07-14
📖 5 min read🧠 Deep dive

Original authors: Avrile Floro (UPHF), Tamara Dhorasoo (UPHF), Soline Pellez (UPHF), Nils Holzenberger

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to solve a mystery in a giant library of French court decisions. Your job is to find the "hidden clues": moments where a judge uses a specific rule from the law book (the Civil Code) to make a decision, but never actually says the rule's name.

It's like a chef making a perfect cake but never listing "flour" in the recipe. You know they used it because the cake tastes like flour, but the text doesn't say so. In the legal world, these are called implicit citations.

The Big Problem: The "Blind Spot"

Usually, lawyers find laws by searching for keywords like "Article 1192." But if a judge applies that article without writing the number, the search engine goes blind. This creates a huge gap in our understanding of how laws are actually used. The authors of this paper wanted to build a computer program (a model) that could act like a super-detective and find these hidden rules.

The Experiment: A Test of Three Experts

To teach the computer, the team created a special test set of 1,015 pairs of text. They asked three legal experts to look at each pair and decide: "Is the judge using this law, or just talking about facts?"

Here is where things got interesting. The experts didn't always agree. In about one-third of the cases (339 of them), the experts argued. One might say, "Yes, that's the law!" while another said, "No, that's just a fact."

The Main Finding: Where Humans Stumble, Computers Stumble Too

The paper's biggest discovery is a bit of a twist. Usually, when experts disagree, we think it's just "noise" or confusion. But the authors found that this disagreement is actually a warning sign.

When the experts couldn't agree on whether a law was being used, the computer models failed almost every time.

  • The best computer team (an "ensemble" of models) got a score of 0.70 overall, which is decent.
  • However, look at the mistakes: Two-thirds of the computer's false alarms (saying "Yes, it's the law!" when it wasn't) happened exactly on those tricky cases where the human experts were fighting.

The paper suggests that these aren't just random errors. They are intrinsic difficulties. The line between "applying a law" and "stating a fact" is blurry, even for humans. When humans are unsure, computers are even more likely to be confidently wrong.

What the Paper Rules Out (The "Not" List)

  • It's not just "noise": The paper argues against the idea that expert disagreement is just messy data that should be ignored. Instead, it's a signal that the case is genuinely hard.
  • It's not a solved problem: The authors do not claim they have built a perfect tool that can replace lawyers. They explicitly state that perfect classification (getting a simple Yes/No right every time) is not currently possible for these tricky cases.
  • It's not just about the computer's brain: The failure isn't because the computer is "dumb." Even when the computer was very confident (giving a high score), it was still wrong on these disputed cases. The paper shows that confidence does not equal accuracy when the humans are arguing.

The Silver Lining: A New Way to Play

So, if the computer can't give a perfect Yes/No answer, is it useless? The authors say no. They suggest a different game: Ranking.

Instead of asking the computer, "Is this the law?" (which it might get wrong), ask it, "Which 200 cases are most likely to be the law?"

  • By letting multiple models vote and ranking the results, they created a "shortlist."
  • When they looked at the top 200 candidates, 76% were actually correct.
  • This means a lawyer could use the tool to scan through thousands of documents and get a manageable list of the most promising "hidden clues" to review manually.

How Sure Are We?

The authors are very careful with their language.

  • They measured the disagreement and the error rates precisely using their dataset of 1,015 cases.
  • They proved that the errors concentrate on disputed cases across ten different models.
  • They suggest (but do not prove) that this pattern exists because of the "open texture" of law—a concept where legal rules have a clear core but fuzzy edges where application is uncertain.
  • They demonstrated that an unsupervised ranking tool (one that doesn't need human labels to work) can reach that 76% precision on the top 200 results.

The Takeaway

The paper concludes that we shouldn't expect computers to perfectly replace human judgment on these fuzzy legal edges. Instead, the best use of AI is to act as a high-powered spotlight. It can't tell you the final verdict, but it can shine a light on the 200 most likely candidates out of a million, letting the human expert do the final, crucial work of deciding what's real and what's just a shadow.

In short: Where experts disagree, models fail. But if you use models to find the "maybe" pile, you can still find the gold.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →