Gated Against One Model, Open to the Next: Option-Only Solvability in Legal Multiple-Choice Benchmarks
This paper reveals that legal multiple-choice benchmarks like UA-JudgeExam are often solvable by models using only answer options without reading the questions due to exploitable option-position biases and distractor patterns, necessitating rigorous gating and debiasing to accurately assess true legal reasoning capabilities.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the world of artificial intelligence, researchers often test how smart a computer is by asking it multiple-choice questions, much like a student taking a standardized exam. The standard way to grade these tests is simple: if the computer picks the right answer, it gets a point. But there is a hidden flaw in this method. A computer does not always need to understand the question to get the answer right. Sometimes, the list of possible answers contains enough clues on its own to reveal the correct choice, even if the question is completely hidden. This happens when the wrong answers are written poorly or when the right answer is the only one that sounds like a real, complete sentence. If a machine can guess the right letter just by looking at the options, the test is not actually measuring intelligence; it is measuring how well the machine can spot a pattern in the choices. This is a serious problem for anyone trying to evaluate legal AI, because if the test is flawed, the results are misleading.
A researcher set out to investigate this issue using a massive bank of real legal questions from Ukraine. These questions were written for judges taking a high-stakes qualification exam, and they come with official, verified answers. The researcher wanted to see if modern AI models could solve these questions without ever seeing the question itself. They took thousands of these items and showed the AI only the four possible answers, hiding the question entirely. They found that the AI could still guess the correct answer far more often than random chance would allow. In fact, one model got the right answer nearly 38 percent of the time just by looking at the options, a score that is far above the 25 percent one would expect from a blind guess. This happened because the correct answers were often self-contained statements of law, while the wrong answers contained subtle errors or sounded unnatural. The AI learned to recognize the "correct-sounding" law without needing to know what the question was asking.
The researcher then tried a common fix: they filtered out the questions that the AI could answer blindly and kept only the ones where the AI failed. They hoped this would leave them with a clean set of questions that truly tested reasoning. However, their experiment revealed a startling truth: this fix does not work. They built their filter using one specific AI model, but when they tested a different, more powerful model on the remaining questions, that new model could still guess the answers blindly with high accuracy. The filter had only removed the questions that were easy for the first model to exploit; it did not remove the questions that were easy for a smarter model to exploit. The study showed that as AI models get better, they get better at spotting these hidden clues in the answer choices, meaning that a test cleaned for one model is instantly useless for a better one.
The researcher also discovered that the problem was not about the subject matter of law itself, but about how the questions were written. They compared the Ukrainian exam questions to a different legal test called LEXam, which uses a different format. In the LEXam test, the options are not full sentences but short pointers like "statement one and three," which refer back to a list in the question. Because these options have no meaning on their own, the AI cannot guess the answer without reading the question. When the researcher tested their AI models on this different format, the models could no longer guess the answers blindly; they scored exactly at the level of random chance. This proved that the flaw was not in the AI's ability to reason, but in the design of the test questions. When the options are full, self-contained statements, the test is vulnerable to exploitation. When the options are just pointers, the test is secure.
The study concludes that simply filtering out bad questions is not a solution, because the definition of a "bad" question changes depending on how smart the AI is. To get a true measure of an AI's ability, researchers must report two numbers: how well the model does with the full question, and how well it does with the question hidden. They must also account for the model's habit of picking certain letters, like always choosing "A," which can artificially inflate scores. The paper suggests that for legal exams and similar professional tests, the format of the answer choices is the most critical factor. If the answers are written as complete propositions, the test will likely be compromised by AI that can read the options alone. If the answers are designed to be pointers, the test remains a valid measure of reasoning. The researcher released their entire dataset and the results of their tests so that others can see exactly which questions leaked information and which did not, providing a clear path forward for building better, more honest evaluations of artificial intelligence.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.