A Two-Phase Stability Study of LLM Judges and Bar Council Examiners on Thai Bar-Exam Free-Form Essays
This study reveals that while LLM judges and human examiners converge on Thai bar-exam essays with clear rubrics, LLMs systematically fail to reproduce a minority human interpretation on ambiguous cases, instead uniformly aligning with the majority human view and thereby masking their inability to capture the full spectrum of expert disagreement.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are running a high-stakes contest to judge the quality of legal essays. You have a very specific rulebook (the "rubric") that tells graders how to score answers. Usually, the rulebook is clear: if the answer is right, give a high score; if the reasoning is bad, give a low score.
But sometimes, the rulebook hits a "gray zone." It doesn't say exactly what to do when a student gets the final answer right but forgets to mention one specific, crucial law citation. This is the "regulation-silent" zone.
This paper is a two-part experiment to see how humans and Artificial Intelligence (AI) judges handle these gray zones.
Part 1: The "Exam" (Phase 1)
First, the researchers put 8 different AI models into the "student seat." They had to write legal essays just like 42 real Thai law students. The AI models were graded blindly by human experts.
- The Result: The AI models performed about as well as the average human student. Some were great, some were average, and some struggled. They were all "in the game."
Part 2: The "Grading" (Phase 2)
Next, the researchers took the same essays and asked two groups to grade them:
- The Human Panel: Three real, certified legal examiners.
- The AI Panel: A massive group of 26 different AI models (including the ones that took the exam earlier, plus many others).
They gave everyone the exact same four things to look at: the question, the rulebook, the "perfect" answer key, and the student's essay.
The Big Discovery: The "One-Way Street"
The researchers found a fascinating split, but only on the tricky questions where the rulebook was silent.
The Human Split:
The three human experts didn't agree.
- Two of them (B and C) said: "The student got the main point right and the reasoning was good enough. Give them a high score (6–8)."
- One of them (A) said: "The student missed a crucial citation. That makes the reasoning 'unusable.' Give them a very low score (1–2)."
- Think of it like a sports referee: Two refs say the player was "in bounds," and one ref says they were "out of bounds." Both are using the same rulebook, but they interpret the gray area differently.
The AI Split (The Asymmetry):
Here is where it gets weird. The researchers expected the 26 AI models to split up, maybe some siding with the strict human (A) and some with the lenient humans (B and C).
- They didn't.
- 22 out of 26 AI models sided with the two lenient humans (B and C). They gave high scores.
- 3 AI models were confused and gave middle scores.
- Zero AI models sided with the strict human (A). Even the one AI that tried to be strict (GPT-5.4 Nano) wasn't consistent enough to truly match the strict human's style.
The Metaphor:
Imagine a group of 26 different cameras taking a photo of a painting.
- The three human art critics look at the painting and disagree: Two say it's a masterpiece, one says it's a failure.
- You ask the 26 cameras to describe the painting.
- The result: All 26 cameras agree with the two critics who called it a masterpiece. None of the cameras agree with the critic who called it a failure. Even though the cameras are different brands, sizes, and prices, they all "see" the painting the same way, missing the perspective of the one strict critic entirely.
Why Does This Matter?
The paper argues that when we build AI systems to grade essays, we usually pick the AI that agrees the most with the "average" human grader.
- Because the majority of humans in this study (2 out of 3) were lenient, the AI learned to be lenient too.
- The AI didn't learn to be "fair" in a balanced way; it learned to be asymmetric. It converged on the majority view and completely ignored the minority view, even though that minority view was a valid, legal interpretation.
The "Dual-Role" Surprise
The researchers also noticed something strange about the AI's personality.
- The AI models that were bad students in Phase 1 (they got low scores on their own essays) actually became the most consistent lenient judges in Phase 2.
- The AI models that were great students in Phase 1 became only "okay" judges.
- The Lesson: Being good at writing the answer doesn't make you good at grading the answer. The skills are totally different.
The Bottom Line
The study shows that AI judges are very stable and consistent with each other, but they have a blind spot. When human experts disagree on a gray-area rule, AI doesn't split the difference or explore all valid opinions. Instead, it overwhelmingly picks the "majority" human opinion and ignores the "minority" one.
If you use AI to grade essays, you aren't just getting a second opinion; you are getting a mirror that reflects the majority view so strongly that it makes the minority view disappear completely.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.