← Latest papers
💬 NLP

Can LLMs Reason in a Legally Meaningful Manner? A Small-scale Study on European Court of Human Rights Cases

This study evaluates a top-tier LLM's ability to perform legally meaningful reasoning on European Court of Human Rights cases, finding that while the model produces structurally complete but substantively shallow analyses and automated evaluators fail to align with human judgment, expert prompting improves reasoning depth without enhancing prediction accuracy.

Original authors: Amogh Raina, Ilias Chalkidis, Daniel Hershcovich, Henrik Palmer Olsen

Published 2026-08-19
📖 5 min read🧠 Deep dive

Original authors: Amogh Raina, Ilias Chalkidis, Daniel Hershcovich, Henrik Palmer Olsen

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the high-stakes world of law, a judge's decision is never just a guess; it is the result of a careful, step-by-step process of reasoning. A judge must look at the facts of a case, check if the law was followed, decide if the government had a good reason for its actions, and finally weigh whether those actions were necessary in a free society. This process is what gives legal rulings their authority and allows people to understand why a decision was made. Today, powerful computer programs known as large language models can read vast amounts of text and answer complex questions, leading many to wonder if these machines can also perform this kind of deep legal reasoning. The question is not just whether a computer can predict the outcome of a court case, but whether it can explain its thinking in a way that a human lawyer would recognize as sound and meaningful.

A team of researchers set out to test this idea using the European Court of Human Rights, a real court that hears cases about whether countries have violated the rights of their citizens. They focused on a specific group of cases involving freedom of expression, where the court must decide if a government's restriction on speech was justified. The researchers took thirty recent cases from the court's official records and asked a top-tier computer model to act as a judge. They gave the model the facts of each case and asked it to produce a legal assessment, explaining its reasoning step-by-step before predicting the final verdict. To see how well the model did, they compared its explanations against the actual reasoning used by the real human judges in those cases. They also tested the model under different conditions: once with no special instructions, once with a detailed guide written by a legal expert, and once with the official guidebook for that specific law, asking the model to figure out the reasoning steps on its own.

The results revealed a significant gap between what the computer could do and what is required for true legal reasoning. The model was surprisingly good at following the basic structure of a legal argument. In almost every case, it produced an answer that looked like a legal judgment, with distinct paragraphs for each step of the analysis. It correctly identified that a restriction on speech had occurred, checked if there was a law behind it, and considered if the government had a valid goal. However, when the researchers looked closer at the quality of the thinking, they found the model's reasoning was often shallow. While the structure was there, the substance was missing. The model frequently offered vague conclusions instead of firm decisions, relied too heavily on findings from lower courts that had been overturned, and failed to grasp the subtle, practical balancing acts that real judges perform. For instance, in one case involving a prisoner's right to receive mail, the real court understood that reviewing thousands of photocopies for security risks was an unreasonable burden on prison staff, a nuance the model missed entirely.

Perhaps the most surprising finding was that the quality of the model's reasoning had almost nothing to do with whether it guessed the correct outcome. The model predicted the right answer about eighty percent of the time, but this accuracy was largely because it simply guessed that a violation had occurred, which happened to be the most common result in the cases they studied. When the model got the answer right, it was often just as shallow in its reasoning as when it got the answer wrong. This suggests that the model is not truly "thinking" through the legal problems; it is recognizing patterns and guessing the most likely result based on the majority of past cases. Even when the researchers gave the model a detailed, expert-written guide on how to reason, the model produced more comprehensive explanations, but this did not make its predictions any more accurate.

The study also examined how to measure the quality of these computer-generated legal arguments. The researchers had human law students and other computer models act as judges to grade the reasoning. They found that while the computer judges were consistent with each other, they did not agree well with the human experts. The computer judges tended to be more lenient and missed the deeper flaws in the reasoning that the humans caught. This indicates that using another computer program to grade the work of a first computer program is not a reliable way to ensure quality. The human experts found that the model's reasoning, while structurally complete, lacked the depth and certainty of a real legal judgment. The model would often hedge its bets or fail to commit to a clear conclusion, whereas a real judge must be decisive.

Ultimately, the research suggests that while these advanced computer models can mimic the outward appearance of legal reasoning, they do not yet possess the ability to reason in a legally meaningful way. They can produce text that looks like a court opinion, but the logic inside is often superficial and disconnected from the complex realities of the law. The study warns that relying on these models for legal decisions, or trusting automated systems to evaluate their own work, is dangerous. The accuracy of a prediction should not be mistaken for the quality of the reasoning behind it. Until these systems can demonstrate a genuine understanding of legal principles rather than just pattern matching, they should remain tools to assist human lawyers, not replace the human judgment that is essential to the justice system.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →