Knowing the Form, Not the Function: Automatically Auditing Answer--Authority Decoupling in Legal Benchmarks
This paper proposes a joint evaluation framework for legal benchmarks that simultaneously assesses answer correctness and statutory authority grounding, revealing that models frequently decouple these dimensions and that relying solely on answer accuracy leads to misleading performance metrics.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are in a giant, high-tech library where robots are learning to be lawyers. In this world, there's a special game called a "benchmark." It's like a final exam where a robot reads a tricky legal question and picks the right answer. For a long time, the teachers (the scientists) only cared if the robot got the final answer right. If the robot said "The correct choice is B," it got a gold star. But here's the catch: sometimes, the robot might say "B" but give a completely made-up or wrong reason for why it's B, like citing a law that doesn't exist or pointing to the wrong rulebook.
This paper dives into a specific corner of that library: the world of Legal Benchmarks. Think of these benchmarks as the standardized tests for AI lawyers. The key idea the paper explores is Authority Grounding. In plain English, this means checking if the robot's answer is actually supported by the real, official laws it claims to be using. It's the difference between a student guessing the right answer on a math test versus actually showing the correct formula. The author cares about this because if we only grade the final answer, we might think a robot is a brilliant lawyer when it's actually just a lucky guesser who happens to sound confident.
The researchers decided to play a game with four different AI models (think of them as four different robot students) using real questions from the Taiwan Bar Exam. They asked the robots to solve legal puzzles and explain their reasoning, but they didn't explicitly tell them, "Hey, you must quote the specific law article." Surprisingly, the robots started quoting laws on their own, just like a real lawyer would.
Here is the big surprise the paper found: The robots often got the answer right but the law wrong, and sometimes they got the answer wrong but the law right. It's like a student getting the right answer on a history test but citing the wrong century as their source. The paper calls this a "double dissociation," which is a fancy way of saying the two skills—getting the answer right and finding the right law—can completely separate from each other.
In the criminal law section of the exam, the paper measured exactly how often this happened. They found that between 24.0% and 42.4% of the time, a robot gave a correct answer but missed the correct law entirely. Even stranger, between 15.2% and 21.7% of the time, a robot gave the wrong answer but correctly cited the right law. This proves that just because a robot gets the final answer right, it doesn't mean it actually understands the legal rules behind it.
The author also tried a little experiment. They told the robots, "If you aren't sure of the exact law number, you don't have to say it." When they did this, the robots stopped quoting laws as much, but their ability to get the final answer right barely changed. This suggests that the robots weren't really using the laws to figure out the answer; they were just adding the laws in as decoration after they had already guessed the answer.
So, what's the takeaway? The paper argues that we need to change how we grade these AI lawyers. We can't just look at the final answer anymore. We have to check the "receipts"—the specific laws they cite. If we don't, we might be celebrating a robot that is actually just hallucinating legal facts while getting lucky with the right answer. The author suggests that future tests should score both the answer and the law citation together, because right now, the current system is missing a huge chunk of the truth.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.