LLM Judges Have Dark Current: A Psychometric Datasheet for LLM-as-a-Judge Evaluation
This paper introduces a "Judge Datasheet" protocol that treats LLM-as-a-judge systems as measurement instruments rather than simple scoring devices, proposing a psychometric framework to quantify specific biases like "dark current" and positional preference to ensure reliable evaluation before making downstream claims.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are hiring a team of art critics to judge a painting contest. You want to know who is the best artist, so you ask these critics to compare two paintings and say which one is better.
This paper argues that we have been treating these "AI Critics" (LLM Judges) too simply. We usually just ask them, "Who won?" and report a single number, like "90% accuracy." The authors say this is like buying a thermometer without checking if it's broken, if it reacts to the wind, or if it gives a temperature reading even when there is no heat.
Here is the paper's core message, broken down with simple analogies:
1. The "Dark Current" Problem (The Phantom Signal)
In physics, "dark current" is when an electronic sensor gives a reading even when there is absolutely no light hitting it.
- The Paper's Finding: The authors tested AI judges by giving them two identical answers (or even empty answers). A good judge should say, "These are the same, I can't pick a winner."
- The Reality: Some judges (like the Llama-3.1-8B model) kept picking a winner anyway, even when the answers were identical. They were "hallucinating" a preference where none existed. This is their "Dark Current."
2. The "Position Bias" (The Seat Preference)
Imagine a judge who always picks the person sitting in a specific chair (e.g., always the left chair), no matter who is actually sitting there.
- The Paper's Finding: The authors tested this by swapping the order of the answers. If the judge picks "Answer A" when it's first, but then picks "Answer B" (which is actually the same as A) when it's first, they aren't judging the content; they are just picking a seat.
- The Reality: One of the judges (Llama-3.1-8B) was almost entirely driven by this "seat preference." It didn't care about the quality; it just wanted to pick the same position regardless of the content.
3. The "Datasheet" (The ID Card for Judges)
Just as you wouldn't buy a car without a spec sheet telling you its horsepower, fuel efficiency, and safety rating, the authors say we shouldn't use an AI judge without a "Judge Datasheet."
This datasheet measures five specific things:
- Dark Current: Does it make up answers when there is no signal?
- Stable Sensitivity: Does it give a consistent verdict on two answers of the same quality when their wording or presentation order changes? (It measures consistency on surface-form variation, not quality detection.)
- Positional Bias: Does it cheat by picking the same slot/position regardless of the actual content after the order is swapped?
- Target Sensitivity: Can it tell the difference between a "good" answer and a "great" answer? (This is the property that detects real/intended quality differences.)
- The "Tie" Button: How strict is it about calling a tie?
4. The Three Judges (The Case Study)
The authors tested three different AI models to see what their "Datasheets" looked like:
- Judge A (Llama-3.1-8B): This judge appears poorly calibrated for same-quality comparisons in this controlled setting. It has high "Dark Current" (it picks winners even when answers are identical) and is almost entirely driven by "Positional False Preference" (it picks the same slot regardless of content). It is less reliable for comparing similar-quality answers, though it might still be useful for pipeline debugging or for comparisons involving larger quality differences.
- Judge B (Qwen2.5-14B): This judge is mixed. It doesn't have "Dark Current" (it stays quiet when there's no signal), and it is very good at spotting big differences in quality. However, when the answers are very similar (specifically same-quality pairs with surface-form variation), its behavior is mixed: sometimes it responds consistently to the surface-form variation, and sometimes it picks based on the order they were shown.
- Judge C (Qwen2.5-32B): This is the cleanest judge. It has no "Dark Current," LOW positional false preference (not zero), and it is very good at spotting real quality differences. However, it is a bit "conservative"—it prefers to say "It's a tie" rather than guess when the difference is very small.
5. The "Strict Tie" Experiment
The authors tried a trick: they told the "cleanest" judge (Qwen2.5-32B), "Be stricter! Only pick a winner if you are 100% sure. Otherwise, call it a tie."
- The Result: This successfully eliminated false preferences on the tested same-quality pairs (pairs with surface-form variation).
- The Catch: It also made the judge miss some real but very small differences. It turned "I think this one is slightly better" into "I'm not sure, it's a tie."
- The Lesson: You can change the judge's "strictness" (the criterion) by changing the instructions, but you cannot magically make the judge smarter or more sensitive just by asking nicely. Note: This experiment did not remeasure the "Dark Current" on identical/empty inputs under the strict prompt.
The Bottom Line
The paper does not claim that one of these judges is the "best" for all human tasks, nor does it prove any specific theory about how AI works.
Instead, it claims that before we trust an AI to judge other AIs, we must first measure the judge itself. We need to know if it has "Dark Current," if it's biased by position, and how strict it is. Without this "Datasheet," any score we get from an AI judge is just a number with no context, potentially hiding serious flaws.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.