JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems
The paper introduces JudgeSense, a benchmark designed to quantify the stability of LLM-as-a-judge systems by measuring how much their verdicts change when prompts are paraphrased, revealing significant inconsistencies in task performance and inherent biases across different models.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The "Moody Judge" Problem: Why Your AI Evaluator Might Be Unreliable
Imagine you are entering a cooking competition. You present a delicious lasagna to a judge.
The first time you ask, "How would you rate this lasagna on a scale of 1 to 10?", the judge gives you a 9.
The next day, you ask the exact same judge, but you phrase it slightly differently: "On a scale of 1 to 10, how tasty is this lasagna?" Suddenly, the judge gives you a 6.
The lasagna hasn't changed. The judge hasn't changed. Even the scale is the same. But because you changed a few words, the verdict flipped. In the world of Artificial Intelligence, this is a massive problem, and a new research paper called JudgeSense has just put a spotlight on it.
The Core Concept: The "Judge Sensitivity Score" (JSS)
Nowadays, instead of humans grading AI models, we use "AI Judges" (like GPT-4 or Claude) to grade other AIs. We assume these AI judges are objective, like a professional referee.
However, the researchers behind JudgeSense discovered that many AI judges are actually quite "moody." They are hyper-sensitive to how a question is phrased. If you change "Is this correct?" to "Is this accurate?", the AI might change its mind.
To measure this "moodiness," they created the Judge Sensitivity Score (JSS):
- A score of 1.0 is like a rock-solid judge: No matter how you ask the question, the answer is always the same.
- A low score is like a judge who is easily swayed by the wind: A tiny change in your wording causes them to flip-flop on their decision.
The Four Big Discoveries
The researchers tested nine different AI models across four different "tasks" (like checking facts or judging how well a summary flows). Here is what they found:
1. The "Coherence" Rollercoaster (The Moodiest Task)
When asking AIs to judge how well a story "flows" (coherence), the results were all over the place. Some models, like Claude, were incredibly stable (almost a perfect 1.0). But others, like Gemini, were wildly inconsistent. It turns out that judging "quality" is much harder for an AI than just checking a fact, making it very easy for the wording to trip them up.
2. The "Polarity Trap" (The Confused Student)
The researchers found a weird glitch in how AIs handle facts. If you ask, "Is this true? Answer YES or NO," the AI might say YES. But if you ask, "Does this contain errors? Answer YES or NO," the AI might get confused and say YES again—even though "YES" now means the opposite! This isn't because the AI is "dumb," but because the way the question is flipped (the "polarity") confuses its logic.
3. Size Doesn't Equal Stability (The "Big is Not Always Better" Rule)
In the AI world, bigger models (with more "brain cells" or parameters) are usually assumed to be smarter. But JudgeSense proved that bigger isn't always more consistent. A smaller, more focused model can sometimes be a much more reliable judge than a massive, expensive "frontier" model that gets distracted by fancy wording.
4. The "Always-A" Glitch (The Lazy Referee)
In some tests, the AI judges became "lazy." When asked to choose between two options (Option A or Option B), most of them simply picked Option A every single time, regardless of what the content actually said. It’s like a referee who always awards the penalty to the home team just because they are standing on the left side of the field. This is called "position bias," and it makes the judge useless for making real comparisons.
The Takeaway: How to Use AI Judges Safely
If you are a developer using AI to grade your work, the paper offers three pieces of "survival advice":
- Don't trust a single question: Because AIs are sensitive to wording, don't rely on one prompt. If the answer changes when you rephrase the question, your judge is unreliable.
- Watch your "Yes/No" logic: Be very careful when asking questions that flip the meaning (like "Is this wrong?" vs "Is this right?").
- Pick for stability, not just "smartness": When choosing an AI judge, don't just pick the most famous or biggest one. Pick the one that proves it can give the same answer twice.
In short: JudgeSense reminds us that even in the age of super-intelligent AI, the way you ask a question is just as important as the answer you get.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.