Evidence Lock Before Commitment: A Frozen Interface Degrades LLM-as-Judge Evaluation
This paper demonstrates that "evidence locking"—a workflow where LLM judges extract and persist evidence in a separate call before making a final verdict—significantly degrades evaluation performance by reducing alignment with human preferences and increasing answer-order inconsistency compared to structured one-call judging.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to solve a mystery, but instead of looking at the crime scene yourself, you are forced to read a police report written by someone else. You have to decide who the culprit is based entirely on that written summary. This is the world of "AI Judges." In the field of artificial intelligence, we often ask smart computer models to act as referees, deciding which of two answers is better. Usually, these AI judges look at the original answers and the question all at once, making a quick call. But recently, a new idea became popular: "Let's make the AI write down its reasons first!" The hope was that if the AI had to explain its thinking and gather evidence before picking a winner, it would be fairer, more logical, and harder to trick. It's like asking a student to show their work on a math test before giving them the grade. But what if that written "work" isn't actually enough to solve the problem? What if the act of freezing that written note and throwing away the original answers actually makes the AI worse at its job?
This paper, titled "Evidence Lock Before Commitment," dives into exactly that question. The researcher, Divyansh Singh from the University of Florida, wanted to test a specific workflow that many people were starting to use. They asked: If we force an AI to write down its evidence in one step, save that note, and then give only that note to a second step to make the final decision (without letting it see the original answers again), does it still make good choices? They tested this with two very smart AI models, Claude Sonnet 4.5 and GPT-5, across 24,000 different judging scenarios.
The results were a bit of a shock. The researcher found that this "Evidence Lock" method actually made the AI judges worse. When the AI was forced to rely only on its frozen notes, it agreed with human preferences 4 to 6 percentage points less often than when it could just look at the original answers. It also became much more confused by the order in which answers were presented, becoming inconsistent 8 to 10 percentage points more often. Interestingly, simply asking the AI to write down evidence while it was still looking at the original answers (a "structured one-call" method) worked just fine. The problem wasn't the writing; the problem was the "locking." By freezing the evidence and cutting off access to the source, the AI lost the ability to see the full picture. The paper suggests that while keeping a written record is great for auditing and checking work later, it shouldn't replace the original source material when the final decision is being made. The "frozen interface" degrades the quality of the verdict, proving that sometimes, you really do need to see the crime scene yourself, not just read the report.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.