← Latest papers
🤖 AI

Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks

This paper introduces Magis-Bench, a novel benchmark derived from Brazilian judicial examinations that evaluates 23 state-of-the-art LLMs on magistrate-level legal reasoning and writing tasks, revealing that even top-performing models struggle to achieve high scores in producing reasoned judicial decisions.

Original authors: Ramon Pires, Thales Sales Almeida, Celio Larcher Junior, Giovana Bonás, Hugo Abonizio, Marcos Piau, Roseval Malaquias Junior, Thiago Laitz, Rodrigo Nogueira

Published 2026-05-12
📖 3 min read☕ Coffee break read

Original authors: Ramon Pires, Thales Sales Almeida, Celio Larcher Junior, Giovana Bonás, Hugo Abonizio, Marcos Piau, Roseval Malaquias Junior, Thiago Laitz, Rodrigo Nogueira

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a high-stakes cooking competition. Usually, when we test AI, we ask it to cook a dish (write a legal argument or a contract). But in a real legal system, the most important job isn't just cooking; it's being the judge who tastes the dish, decides if it's good, and explains why it passes or fails.

This paper, Magis-Bench, is about testing AI on that "judge" role.

The Problem: Can AI Be a Judge?

Most AI tests in law ask the computer to be a lawyer: "Write a defense for this client." But a real judge has to do something harder: listen to both sides, ignore their own feelings, apply strict rules, and write a final decision that stands up to scrutiny.

The authors wanted to know: Can current AI models act like a real judge?

The Test: The "Magistrate" Exam

To test this, the researchers didn't make up fake questions. They grabbed 74 real questions from actual, highly competitive Brazilian exams used to hire real judges between 2023 and 2025.

Think of these exams as the "Bar Exam" for judges. They include:

  • Essay Questions: "Here is a messy legal situation; explain the law."
  • Drafting Tasks: "Here are the facts; write the full, official court sentence (the final verdict)."

Crucially, every question came with an official answer key (a rubric). This is like a teacher's grading sheet that says exactly how many points you get for mentioning a specific law or following a specific structure.

The Method: AI vs. AI

Since hiring 23 human judges to grade 74 essays for 23 different AI models would be incredibly expensive and slow, the researchers used a clever trick: They let AI grade AI.

  1. The Contestants: They picked 23 of the smartest AI models available (from companies like Google, OpenAI, Anthropic, and others).
  2. The Judges: They used four other, very powerful AI models to act as the graders.
  3. The Rules: The "Judge AIs" were blind to who wrote the answers. They just looked at the question, the official answer key, and the AI's response, then gave it a score from 0 to 10.

The Results: The AI "Judges" Agreed

Here is what happened:

  • The Judges Got Along: The four AI judges agreed with each other almost perfectly (98.4% agreement). This is like having four food critics who all agree on exactly which restaurant is the best. This gave the researchers confidence that the scores were fair and not just random.
  • The Winners: The top performers were Google's Gemini-3-Pro-Preview (score: 6.97/10), followed by Gemini-3-Flash and Claude-4.5-Opus.
  • The Reality Check: Even the "winner" only got 6.97 out of 10. That is less than 70%.

What This Means

The paper concludes that while AI is getting better at legal tasks, acting like a judge is still very hard for them.

  • The Gap: Current AI models are like students who can recite the textbook but struggle to apply the rules to a messy, real-life situation with the nuance and authority of a human judge.
  • The Benchmark: The researchers released all the questions, the answers, and the grading code so other scientists can keep testing and improving these models.

In short: We asked AI to take a real judge's exam. The smartest ones passed with a "C" or a "B," proving that while they are getting smarter, they aren't ready to replace human judges in the courtroom just yet.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →