CAREBench: Evaluating LLMs' Emotion Understanding by Assessing Cognitive Appraisal Reasoning
The paper introduces CAREBench, a novel benchmark grounded in appraisal theory that evaluates LLMs' emotion understanding through cognitive reasoning chains rather than simple label prediction, revealing that current models struggle with appraisal reasoning and positive emotion recognition despite matching human performance on certain tasks.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a robot how to understand human feelings. For a long time, we tested these robots by asking them a simple question: "What is this person feeling?" and expecting a single answer like "Sad" or "Happy."
The authors of this paper, CAREBench, argue that this is like testing a chef by only asking, "Is the soup salty?" without ever asking why it tastes that way or how they made it. A robot might guess "salty" correctly just by memorizing patterns, without actually understanding the cooking process.
Here is a simple breakdown of what they did and what they found, using everyday analogies.
1. The Problem: The "Black Box" of Emotion
Current tests for AI emotion understanding are too shallow. They treat emotions like a vending machine: you put in a story, and the machine spits out a label (e.g., "Angry"). But real human emotions are more like a complex recipe. They depend on how a person interprets a situation.
- The Old Way: "The car broke down." -> AI says: "Frustrated."
- The Real Way: "The car broke down" -> Interpretation: "I'm late for work and I can't fix this myself" -> Result: "Frustrated."
The paper argues that if an AI skips the "Interpretation" step, it doesn't truly understand the emotion; it's just guessing based on keywords.
2. The Solution: The "CAREBench" Recipe Book
To fix this, the researchers built a new test called CAREBench. Instead of just asking for the final emotion, they created a dataset that forces the AI to show its work, step-by-step.
They collected 1,000 real-life stories from people about emotional events. Then, they asked humans to annotate these stories in four layers, creating a complete "chain of thought":
- The Story: What happened?
- The Reasoning: Why does this matter? (e.g., "This ruins my plan.")
- The Rating: How intense is this feeling? (e.g., "I feel 80% in control.")
- The Emotion: What is the final feeling? (e.g., "Anxious.")
Crucially, they did this from two perspectives:
- First-Person: The person who lived the story.
- Third-Person: An observer reading the story.
This is like having a detective (the AI) try to solve a crime by looking at the suspect's diary (First-Person) and also by interviewing a witness who saw the event (Third-Person).
3. The Test: Six AI Models in the Kitchen
The researchers tested six different Large Language Models (LLMs)—some famous ones like GPT and Claude, and some specialized ones trained on psychology data. They asked the models to perform every step of the chain, not just the final emotion guess.
4. The Results: The AI is a Good Guessers, but a Bad Chef
Here is what they discovered, using simple metaphors:
- The "Good News" (Surface Level): When the AI just had to guess the final emotion (like "Is this person happy or sad?"), the smartest models performed as well as, or even better than, human observers. They are great at spotting the obvious signs.
- The "Bad News" (Deep Understanding): When asked to explain why the person felt that way (the reasoning step), the AI struggled. It often gave vague or generic answers. It's like a student who gets the right answer on a math test but can't show the work.
- The "Positive Emotion" Blind Spot: The AI was surprisingly bad at identifying positive emotions (like "hopeful" or "grateful"). It was much better at spotting negative ones (like "angry" or "sad"). It's as if the robot is tuned to hear alarms but misses the sound of laughter.
- The "Chain Reaction" Failure: The researchers tried to help the AI by giving it the "reasoning" text first, hoping it would use that to get the emotion right.
- Result: When humans provided the reasoning, the AI got better.
- Result: When the AI tried to write its own reasoning first, it actually got worse. It tended to over-dramatize its own thoughts, which confused its final answer.
- The "Human Variety" Gap: Humans don't all react the same way. If you read a sad story, one person might feel "sad," another "angry," and another "sympathetic." The AI models failed to capture this variety. They tended to give a single, average answer, missing the fact that different people feel differently about the same event.
5. The Big Takeaway
The paper concludes that current AI models are "surface-level" emotion detectors. They are excellent at pattern matching to guess the final label, but they haven't truly internalized the complex mental process of how humans generate emotions.
If we only test AI on whether they can guess the right emotion label, we are overestimating their intelligence. We need to test them on their ability to reason through the "why" and "how," just like we test a student on their understanding of the subject, not just their ability to memorize the answer key.
In short: The AI can tell you what someone is feeling, but it still doesn't fully understand why they feel that way, nor does it grasp that different people might feel differently about the same situation.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.