Expressing Social Emotions: Misalignment Between LLMs and Human Cultural Emotion Norms
This paper reveals that current large language models systematically misalign with human cultural norms in expressing social emotions, exhibiting a deterministic bias toward engaging emotions and failing to capture the cultural diversity and variability observed in human interactions.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a super-smart robot friend who has read almost every book, article, and website on the internet. You call this robot "LLM" (Large Language Model). You ask it to pretend to be a person from different cultures—say, a person from the United States or a person from Mexico or Chile—and tell you how they would feel in different situations.
The big question this paper asks is: Is this robot friend actually good at pretending to be a human from a specific culture, or is it just guessing based on stereotypes?
Here is the story of what the researchers found, explained simply.
1. The Two Types of Feelings: "Hugging" vs. "High-Fiving"
To understand the experiment, we need to know about two types of emotions the researchers studied:
- Engaging Emotions (The "Hugs"): These are feelings that bring people together, like guilt, shame, or feeling friendly. They say, "We are all connected."
- Disengaging Emotions (The "High-Fives"): These are feelings that highlight individuality, like pride, anger, or self-esteem. They say, "I am my own person."
The Human Reality:
- Latin Americans (Collectivist culture): In real life, people from Mexico and Chile tend to "hug" more. They express engaging emotions (like feeling guilty if they let a friend down) more often than "high-fiving" emotions.
- European Americans (Individualist culture): In real life, people from the US tend to "high-five" more. They express disengaging emotions (like feeling proud of their own success) more often than "hugging" emotions.
2. The Robot's Big Mistake
The researchers asked six different super-smart robots to pretend to be these people. They expected the robots to mimic the real human patterns.
What happened?
The robots failed spectacularly.
- The "One-Size-Fits-All" Problem: No matter which culture the robot was pretending to be, it acted the same way. It loved "hugging" emotions (engaging) way more than "high-fiving" emotions (disengaging).
- The American Paradox: This was the strangest part. Even when the robot was pretending to be a European American (a culture that usually loves "high-fiving" and individualism), the robot still acted like it loved "hugging" and connection. It completely missed the mark on the culture it is supposedly trained on the most!
The Analogy:
Imagine asking a robot to pretend to be a cowboy and a ballet dancer.
- Real cowboys might be more independent and rugged.
- Real ballet dancers might be more collaborative and graceful.
- But this robot, when asked to be a cowboy, started doing ballet moves. When asked to be a ballet dancer, it also did ballet moves. It couldn't switch costumes; it just kept doing what it thought was "nice" (being polite and connected), ignoring the actual character it was supposed to play.
3. The "Boring Robot" Problem
The researchers also looked at how the robots answered.
- Humans are messy: If you ask 100 real people how angry they are, you get a wide range of answers. Some say "a little," some say "a lot," some say "not at all." It's a colorful rainbow of opinions.
- Robots are boring: The robots were incredibly repetitive. If you asked the same question 190 times, the robot would give almost the exact same answer every time. It was like a broken record stuck on one note.
- The Result: The robots didn't capture the diversity of human feelings. They were too predictable and too safe.
4. Why Did This Happen? (The "Language" Twist)
The researchers tried to fix the robot by changing the settings:
- Turning up the "Randomness" (Temperature): They tried making the robot more random and creative. It didn't help much. The robot just got a little more chatty, but it still gave the same wrong answers.
- Changing the Language: They asked the robot to pretend to be Mexican, but they gave the instructions in English instead of Spanish.
- The Surprise: The robot actually did a better job when asked in English!
- Why? The robot was trained mostly on English text from the internet. Even though it was pretending to be a Spanish speaker, its "brain" understood the cultural concepts better when they were described in English. It's like a student who learned about French culture by reading English travel blogs; they might know more about the culture in English than in French.
5. Why Should We Care?
This isn't just a funny mistake; it's dangerous.
- Mental Health: If you use an AI chatbot for therapy, and the bot is supposed to understand your culture, but it keeps acting like a "nice, polite, connected" person when you are actually an angry, independent American, it might give you bad advice. It might tell you to "apologize" when you should be "standing your ground."
- Fake People: Scientists are starting to use robots to simulate human surveys. This paper says, "Stop doing that!" The robots are too boring and too wrong to replace real humans in studying how different cultures feel.
The Bottom Line
Current AI is like a very well-read, very polite tourist who has read a lot of travel guides but has never actually lived in the places they are describing. They think everyone wants to be polite and connected (the "hug"), but they completely miss the unique, sometimes messy, sometimes proud, sometimes angry ways that real humans actually express themselves in different cultures.
Until we fix this, we have to be careful about trusting AI to understand our feelings, especially when those feelings depend on where we come from.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.