CLASH: Evaluating Language Models on Judging High-Stakes Dilemmas from Multiple Perspectives
The paper introduces CLASH, a dataset of high-stakes dilemmas with diverse character perspectives, to reveal that even advanced language models struggle with value-based decision-making, exhibiting specific failure patterns like early commitment and limited understanding of value shifts despite reasonable predictions of psychological discomfort.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a robot how to make tough life choices. Most previous tests asked the robot simple questions like, "Is it okay to jaywalk if you're late for work?" But in the real world, the hardest choices aren't simple; they are high-stakes dilemmas where every option hurts someone, and values clash like two storms colliding.
This paper introduces CLASH (Character perspective-based LLM Assessments in Situations with High-stakes), a new "exam" designed to see if AI can handle these messy, high-pressure situations.
Here is a breakdown of what they did and what they found, using simple analogies:
1. The Exam: CLASH
Think of previous AI tests as a multiple-choice quiz with short, clean sentences. CLASH is more like a drama series.
- The Stories: Instead of one-liners, the dataset contains 345 long, detailed stories about life-or-death or life-changing situations (like a doctor deciding on a treatment, or a bank teller making a costly mistake).
- The Characters: For every story, the AI has to answer from the perspective of different "characters." Some characters care only about safety; others care only about fairness. Some characters change their minds halfway through the story.
- The Goal: The exam asks: "If you were Character A, would you do this? Would it make you feel sick to your stomach? Did your mind change from yesterday to today?"
2. The Big Findings: Where the AI Stumbles
The researchers tested 14 different AI models (including the newest, smartest ones like GPT-5 and Claude-4) on this exam. Here is what happened:
A. The "Can't Decide" Problem (Ambivalence)
- The Metaphor: Imagine standing in front of two doors. One leads to a reward, the other to a trap. A human might say, "I'm torn; I can't pick."
- The Result: The AI models are terrible at admitting they are torn. Even the smartest models forced a "Yes" or "No" answer when the situation clearly required a "Maybe." They struggled to say, "I don't know," even when the right answer was "I'm conflicted." In fact, the best model only got this right about 51% of the time.
B. The "Gut Feeling" Problem (Discomfort)
- The Metaphor: Sometimes you know what you should do, but it feels wrong in your gut. Like a parent who has to fire a friend to save the company.
- The Result: The AI is actually pretty good at spotting when a character should feel uncomfortable. If the story says, "Character A values honesty but has to lie," the AI correctly guesses, "Yes, Character A feels bad about this."
C. The "Memory" Problem (Value Shifts)
- The Metaphor: Imagine a character who loves pizza, but then gets a stomach ache and decides they hate pizza now.
- The Result: The AI models are terrible at updating their "minds." When a story says, "Character A used to value X, but now values Y," the AI often forgets the change and answers based on the old belief. It's like a student who studied for a test last year but forgot they learned a new subject this year. The accuracy dropped by nearly 43 points when values shifted!
3. How the AI "Thinks" (Reasoning Chains)
The researchers peeked inside the AI's "brain" to see how it solved these problems.
- The Old Trick: In math or chess, smart AIs use a strategy called "backtracking" (checking their work, looking back at the rules). This works great for math.
- The New Failure: In moral dilemmas, this backtracking didn't help. Instead, the AI developed two bad habits:
- Early Commitment: The AI picks a side too quickly, like a judge banging a gavel before hearing the full story.
- Overcommitment: Once it picks a side, it stubbornly sticks to it, ignoring evidence that suggests the other side might be right. It's like a lawyer who refuses to listen to the opposing argument once they've started their closing statement.
4. The "Steering" Test
The researchers also tested how easily they could "steer" the AI toward a specific value (like Safety vs. Self-Esteem).
- The Finding: If an AI already really likes "Safety," it is very hard to make it choose "Self-Esteem" instead. The more the AI naturally prefers one value, the harder it is to steer it toward the opposite one.
- The Perspective Trick: The AI listened better when asked, "What would Character A do?" (Third-person) rather than "What would you do?" (First-person). However, for the value of "Safety," the AI actually listened better when asked in the first person, likely because safety is a deeply ingrained instinct for the model.
Summary
The paper concludes that while AI is getting smarter at math and games, it is still very clumsy at navigating the messy, emotional, and changing world of human values. It struggles to admit when it's confused, it forgets when people change their minds, and it often gets stuck in its own stubborn logic.
Important Note: The authors emphasize that this is a diagnostic tool. They are not saying these AI models are ready to be doctors or judges. In fact, they warn that a high score on this test doesn't mean an AI is safe to use in real life; it just means we now have a better way to see where the AI is failing.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.