← Latest papers
💬 NLP

GradeLegal: Automated Grading for German Legal Cases

This paper addresses the bottleneck in grading German legal exams by systematically evaluating 27 large language models, finding that reasoning-oriented models combined with sample solutions and rubrics can effectively approximate expert grading in public law (though less so in criminal law) and that ensembling strategies can further improve reliability.

Original authors: Abdullah Al Zubaer, Lorenz Wendlinger, Simon Alexander Nonn, Michael Granitzer, Jelena Mitrovic

Published 2026-05-21
📖 4 min read☕ Coffee break read

Original authors: Abdullah Al Zubaer, Lorenz Wendlinger, Simon Alexander Nonn, Michael Granitzer, Jelena Mitrovic

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a teacher in a very strict school where grading exams is like solving a complex mystery. In Germany, law students take written exams that are incredibly long (sometimes 10 to 30 handwritten pages!) and are graded on a scale of 0 to 18. Getting a high score is crucial for their future careers.

The problem? There are too many students and not enough expert teachers to grade them all. It takes weeks or even months to get results, creating a huge bottleneck.

This paper, GradeLegal, asks a simple question: Can Artificial Intelligence (AI) act as a reliable grader for these tough German law exams?

Here is what the researchers found, explained through everyday analogies:

1. The "Blank Canvas" vs. The "Recipe Book"

The researchers tested 27 different AI models (some free and open, some paid and private) to see how well they could grade student answers. They tried different ways of asking the AI to do the job:

  • The "Guess What I'm Thinking" Approach (Task-Agnostic): They just told the AI, "Grade this."
    • Result: The AI got confused. It was like asking a chef to cook a meal without giving them a recipe or ingredients. The AI's grades were almost random, sometimes even worse than a coin flip.
  • The "Show Me the Answer" Approach (Sample Solution): They gave the AI a perfect example of what a top student should write.
    • Result: Better, but still not perfect. The AI knew what the "right" answer looked like but didn't know exactly how to deduct points for small mistakes.
  • The "Rulebook" Approach (Rubric): They gave the AI a strict checklist (a rubric) that said, "If you mention X, give 2 points. If you miss Y, subtract 1 point."
    • Result: This was a game-changer. The AI suddenly understood how to translate a student's messy legal arguments into a specific number.
  • The "Super-Teacher" Approach (Rubric + Sample Solution): They gave the AI both the perfect example and the strict rulebook.
    • Result: This was the winner. When the AI had both, it graded almost as well as a human expert in Public Law (a specific type of law).

2. The "Criminal Law" vs. "Public Law" Puzzle

The researchers tested two types of law exams:

  • Public Law: These exams were like a structured puzzle with a clear path. The AI did very well here, matching human experts when given the right tools.
  • Criminal Law: These exams were like a chaotic maze. The questions were more open-ended, and students could argue their case in many different, complex ways.
    • Result: The AI struggled more here. Even with the best tools, it couldn't quite match the human experts as easily as it did with Public Law. It's like the difference between grading a math test with one right answer (Public Law) versus grading a creative writing contest where there are many valid interpretations (Criminal Law).

3. The "Teamwork" Effect (Ensembling)

The researchers wondered: What if we don't just use one AI, but a team of them?
They tried combining the opinions of three different AI models to create a "super-grader."

  • The Analogy: Imagine asking three different judges to score a gymnast and then taking the middle score.
  • Result: This "team" approach often did better than the single best AI model. It smoothed out the weird mistakes one AI might make. This is exciting because it means a group of free, open-source AIs can sometimes beat the most expensive, paid AI models.

4. The "Reasoning" Factor

They tested "Reasoning" models (AIs that think step-by-step) vs. "Non-Reasoning" models (AIs that just predict the next word).

  • Result: The "Reasoning" models were much better at understanding the deep logic of legal arguments. It's like the difference between a robot that just memorizes a dictionary and a detective who actually investigates the clues.

The Bottom Line

The paper concludes that AI can help grade German law exams, but only if you treat it like a junior assistant, not a magic wand.

  • You must give it a clear rulebook (rubric) and a sample answer.
  • It works best on structured exams (Public Law) and is still learning how to handle the messy complexity of Criminal Law.
  • It is currently best used as a tool for practice and self-testing (helping students see how they might do before the real exam), rather than for making final, life-changing decisions about their grades.

In short: AI is a powerful new teaching assistant, but it needs a very clear set of instructions and a human supervisor to do its best work.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →