Can LLMs Judge Legal Accuracy? Reliability of LLM Evaluators for High-Stakes Insurance QA in a Low-Resource Language
This paper demonstrates that while LLMs can be used to evaluate legal accuracy in low-resource languages like Quebec French, their reliability is insufficient for high-stakes deployment without human oversight, as even the strongest models exhibit moderate agreement, fail to consistently catch dangerous errors, and benefit significantly from step-by-step reasoning and ensemble methods.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the world of artificial intelligence, a new tool has emerged to solve an old problem: how do we know if a computer is telling the truth? For years, developers have relied on human experts to read and grade the answers generated by these smart machines, but this process is slow, expensive, and hard to repeat. To speed things up, researchers began using the smartest artificial intelligence models themselves to act as judges, grading the work of other models. This approach, often called "AI judging," has become popular because it is fast and consistent. However, most of what we know about these digital judges comes from tests done in English on general topics like writing stories or answering trivia. We do not yet know if these judges can be trusted when the stakes are high, the subject is complex law, and the language is a specific regional dialect.
A team of researchers at Université Laval in Canada decided to test these digital judges in the most demanding environment they could imagine. They focused on automobile insurance in the province of Quebec, a place where the rules are written in a unique form of French and governed by a specific legal system that differs from the rest of Canada and Europe. In this system, bodily injury from car accidents is handled by a public, no-fault plan, while damage to property is handled by private insurance companies. A single mistake in explaining these rules could mislead a citizen about their rights or leave a company legally exposed. The researchers gathered 807 real exam questions used to certify insurance professionals, along with the correct answers and expert explanations. They then asked 52 different artificial intelligence models to act as judges, checking whether various answers to these questions were right or wrong.
The results were sobering. Even the most powerful and advanced artificial intelligence models, those considered the best in the world, could not be trusted to judge these legal answers on their own. The best-performing model was only moderately reliable, meaning it still made mistakes often enough that a human would not want to rely on it blindly. More importantly, the researchers found that a model's overall intelligence did not predict how safe it was to use. Some very smart models were surprisingly careless, frequently approving wrong answers as if they were correct. Others were overly cautious, rejecting correct answers just to be safe. There was no simple way to pick a "good" judge based on its general reputation; a model that scored highly on overall quality could still be dangerous because it failed to catch specific, costly errors.
The study also revealed that these digital judges behave differently depending on who is writing the answer. In many other studies, artificial intelligence models tend to be kinder to answers written by other machines, often giving them higher scores than human-written text. But in this strict legal setting, the opposite happened. The judges were actually harsher on answers written by artificial intelligence than on answers written by humans. They were less likely to accept a machine-generated answer as correct, even when the human-written distractors were clearly wrong. This suggests that the bias of these models is not fixed; it changes based on the task and the instructions given. When asked to be precise about the law, the models became stricter, not more lenient.
To find a way forward, the researchers tested different strategies to make the judging process safer. They discovered that giving the judges a reference text to compare against helped, but only for the weaker models. The most effective solution was to use a panel of judges rather than a single one. By having a small group of the best models vote on an answer, and requiring that they all agree before accepting it as correct, the researchers drastically reduced the number of dangerous mistakes. This "conservative" approach meant that if even one judge in the group spotted an error, the answer was rejected. This method lowered the rate of false approvals by more than half compared to using the single best model alone.
The researchers concluded that for high-stakes legal questions, a single artificial intelligence judge is not trustworthy enough to work alone. The risk of a machine approving a wrong legal answer is too great to ignore. Instead, the best approach is to use a small team of smart models to do the initial screening, but to always have a human expert ready to make the final decision on difficult or risky cases. This creates a system where the speed of artificial intelligence is combined with the safety of human oversight. The study serves as a warning that while these tools are powerful, they are not perfect substitutes for human judgment, especially when the rules are complex and the consequences of a mistake are severe.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.