Safety is Contextual, LLM-Judges Are Not: Navigating the Rigid Priors of Evaluators
This paper investigates the limitations of LLMs-as-judges in safety evaluation, revealing that while they can learn from new context, they often rigidly adhere to their internal safety priors rather than adapting to contradictory information or definitions.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have hired a team of super-smart referees (these are the AI models, or "LLM-judges") to watch a massive sports tournament and decide which plays are "fair" and which are "cheating."
The problem is, the rules of the game change depending on where you are playing. In one country, a specific move is a foul; in another, it's a brilliant play. Also, new slang and new types of tricks are invented every day.
This paper asks a simple but scary question: Are these AI referees actually following the rules you just gave them, or are they just playing by their own old, internal rulebook?
The authors found that the referees are surprisingly stubborn. Here is the breakdown using simple analogies:
1. The "Stubborn Referee" Problem (Steerability)
Usually, when you hire a referee, you give them a rulebook. You say, "If someone touches the ball with their hands, it's a foul."
- What the paper found: Even when you explicitly write a new rulebook in front of the AI, it often ignores you. It keeps judging based on the "safety rules" it learned while it was in school (during its training).
- The Analogy: Imagine you hire a referee who grew up playing soccer. You tell him, "Today, we are playing basketball, so using your hands is fine, but kicking the ball is a foul." The referee might still blow the whistle for using hands because, deep down, he thinks, "No, that's a foul! I learned that in soccer!"
- The Twist: The paper found that if you trick the referee by saying, "Let's just call this 'Category A' or 'Category B' instead of 'Safe' or 'Unsafe'," the referee suddenly becomes much better at following your new rules. It's as if the referee only gets confused when the words "Safety" and "Rules" are used, triggering their old training.
2. The "Distracted Student" Problem (Susceptibility)
Sometimes, you give the referee a cheat sheet or a few examples of what to look for right before the game starts.
- What the paper found:
- Old Tricks: If you give the referee examples of "cheating," they mostly ignore them. They stick to what they already know.
- New Knowledge: However, if you give them information about something brand new that they don't know yet (like a new slang word or a news event from next year), they will listen.
- The Analogy: Imagine a student taking a test.
- If you whisper, "Remember, the answer is always 42," and the student already knows the answer is 42, they ignore you.
- But if you whisper, "The answer to this question about a new planet is 42," and the student has never heard of that planet, they will believe you.
- The Catch: The referee only listens to new info if they are unsure about the topic. If they are confident in their old knowledge, they will ignore your new instructions, even if you are right.
3. The "One-Size-Fits-All" Myth
The paper tested 13 different AI referees. They found that:
- Being a "good" referee (getting high scores on standard tests) doesn't mean you are "flexible."
- Being "flexible" (listening to new rules) doesn't mean you are "accurate."
- Some referees are naturally stubborn; others are naturally suggestible. There is no single "best" referee for every job.
The Big Takeaway
The authors conclude that safety is contextual, but AI judges are rigid.
If you are a company trying to use an AI to check if content is safe for a specific culture or a specific new situation, you cannot just assume the AI will follow your instructions. It is likely judging based on its own internal "gut feeling" about what is safe, which might be totally different from your company's policy.
In short: You can't just tell the AI, "Follow these rules." You have to understand that the AI has its own deeply ingrained habits that are very hard to change, unless you trick it into thinking it's playing a completely different game.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.