← Latest papers
💬 NLP

Mining Legal Arguments to Study Judicial Formalism

This study introduces the MADON dataset and a three-stage NLP pipeline to automatically classify judicial reasoning and formalism in Czech Supreme Court decisions, challenging prevailing narratives about formalism in Central and Eastern Europe while demonstrating the potential of legal argument mining for computational legal studies.

Original authors: Tomáš Koref, Lena Held, Mahammad Namazov, Harun Kumru, Yassine Thlija, Ivan Habernal

Published 2026-03-24
📖 5 min read🧠 Deep dive

Original authors: Tomáš Koref, Lena Held, Mahammad Namazov, Harun Kumru, Yassine Thlija, Ivan Habernal

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a giant library where every book is a court decision. For decades, legal scholars have whispered a rumor about the libraries in Central and Eastern Europe (CEE): "The judges here are robots."

The rumor goes like this: These judges don't think about fairness, justice, or the "spirit" of the law. Instead, they are "formalists"—mechanical robots who only look at the exact words on the page. If a law says "file by day 15," and you file on day 16 because your lawyer was in the hospital, a "robot judge" throws your case out. A "human judge" would say, "Wait, let's look at the bigger picture."

The Problem:
Until now, no one could prove if this rumor was true or just gossip. Reading 300,000 court decisions by hand would take a human team of lawyers a lifetime. It's like trying to find a specific needle in a haystack the size of a mountain.

The Solution: The "Legal Detective" AI
This paper introduces a team of AI detectives (using advanced Large Language Models) to solve the mystery. They built a system to read Czech court decisions and answer three big questions:

  1. Can the AI spot where the judge is actually making an argument?
  2. Can the AI tell if the judge is using a "robot" argument (strict text) or a "human" argument (values, fairness, consequences)?
  3. Can the AI decide if the whole decision was "robotic" or "human"?

The Toolkit: How They Did It

1. The "MADON" Dataset (The Training Camp)
Before the AI could go to work, the researchers had to teach it. They gathered 272 real court decisions and hired legal experts to act as "coaches."

  • The Analogy: Imagine teaching a dog to fetch. You don't just throw a ball; you show it the ball, say "fetch," and reward it.
  • The Process: The experts labeled every paragraph of the 272 decisions. They marked which paragraphs contained arguments and what kind of argument it was (e.g., "Linguistic," "Historical," "Practical Consequences"). They also gave each whole decision a "Holistic Label": Formalistic (Robot) or Non-Formalistic (Human).

2. The Three-Stage Pipeline (The Assembly Line)
Instead of throwing the whole text at a giant, expensive AI brain, they built a smart assembly line to save money and make the process clearer:

  • Stage 1 (The Filter): A small, fast AI scans the document and throws away 85% of the text (like boring procedural lists or repetitive phrases) that doesn't contain real arguments. It's like a bouncer at a club checking IDs and letting only the relevant people in.
  • Stage 2 (The Classifier): A powerful AI reads the remaining "argumentative" text and sorts the arguments into 8 different buckets (e.g., "This is about the text of the law," "This is about fairness").
  • Stage 3 (The Verdict): A final calculator looks at the mix of arguments. If the judge used mostly "fairness" and "consequences" arguments, the AI says, "This is a Human Judge." If they only used "text" arguments, it says, "This is a Robot Judge."

The Big Surprise: The Rumor Was Wrong!

When the AI finished reading the data, the results were shocking. The "Robot Judge" narrative was mostly a myth.

The Findings:

  • Text is Rare: The judges rarely relied on the strict, literal wording of the law (only about 5.7% of their arguments). They weren't reading the dictionary; they were thinking.
  • Values Rule: The most common arguments were about Principles, Values, and Teleology (the "purpose" of the law). They were asking, "What is the right thing to do?" and "What are the consequences?"
  • Precedent is King: Instead of ignoring past cases (as the "robot" rumor suggested), these judges relied heavily on Case Law (previous decisions). They were actually very good at following the "human" tradition of looking at how similar cases were decided before.

The Twist:
The study did find a difference between the two highest courts in Czechia:

  • The Supreme Court (older, inherited from the communist era) was slightly more "robotic" in the past.
  • The Supreme Administrative Court (newer, created in 2003) has become increasingly "human" and open-minded over time.

Why This Matters

This study is like a lie detector test for legal history.

  • For Scholars: It proves that you can use AI to study philosophy and judicial behavior at a massive scale, something that was previously impossible.
  • For the Public: It challenges the idea that Eastern European judges are cold, mechanical robots. It shows they are actually engaging in complex, value-based reasoning, just like judges in the US or Western Europe.

In a Nutshell:
The researchers built a high-tech "argument miner" to dig through thousands of court cases. They expected to find a mountain of robotic, text-only decisions. Instead, they found a garden full of judges arguing about fairness, values, and the greater good. The "robot judges" were mostly a myth, and AI helped us finally see the truth.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →