Investigating Counterfactual Unfairness in LLMs towards Identities through Humor
This paper investigates counterfactual unfairness in large language models by analyzing how humor responses change when speaker and target identities are swapped, revealing that models systematically refuse, judge as malicious, and rate as more harmful jokes told by privileged speakers compared to those from marginalized groups.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a very smart, well-read robot named "Comedy Bot." This robot has read almost everything on the internet, so it knows a lot about jokes, culture, and how people talk to each other. But because it learned from the internet, it also absorbed all the messy, unfair, and sometimes hurtful assumptions humans have about each other.
This paper is like a group of detectives giving Comedy Bot a series of "identity swap" tests to see if it treats people fairly. They use humor as their magnifying glass because jokes are tricky: they depend entirely on who is telling the joke, who is hearing it, and the relationship between them.
Here is what they found, explained through simple analogies:
1. The "Double Standard" Doorkeeper (Task 1)
Imagine Comedy Bot is a bouncer at a club. You ask it to tell a joke.
- Scenario A: A wealthy, powerful person asks for a joke making fun of a poor person.
- Scenario B: A poor person asks for a joke making fun of a wealthy person.
The Result: The bouncer slams the door shut on Scenario A 67.5% more often than Scenario B.
- The Metaphor: It's like a security guard who is terrified of a rich person making a joke about a janitor, but is totally fine with the janitor joking about the rich person. The bot has internalized a rule: "If you have power, you can't punch down." But it's so scared of making mistakes that it sometimes blocks the joke even if it's harmless, just because of who asked.
2. The "Suspicious Mind" Detective (Task 2)
Now, the detectives give the bot the exact same joke to analyze, but they change the characters telling it.
- Joke: "Why did the chicken cross the road? To get to the other side." (Let's pretend this is a joke about body image).
- Scenario A: A fit person tells this to a person with a disability.
- Scenario B: A person with a disability tells this to a fit person.
The Result: Even though the joke is identical, the bot thinks the fit person is being malicious and mean, while it thinks the disabled person is just being friendly or making fun of themselves.
- The Metaphor: It's like watching a movie where the villain is always assumed to be the rich guy in the suit, even if he's just saying "Hello." The bot looks at the "uniform" (the identity) of the speaker and decides their intent before they even finish the sentence. It assumes privilege equals bad intentions.
3. The "Social Mirror" (Task 3)
Finally, the bot has to imagine how a listener would react to a joke.
- Scenario A: A boss makes a joke to an employee.
- Scenario B: An employee makes the same joke to a boss.
The Result: The bot predicts that the employee will be much more offended and the boss will be much more sensitive when the boss tells the joke. It gives the boss a "harm score" that is up to 1.5 points higher on a 5-point scale.
- The Metaphor: It's like a mirror that distorts reality. When a powerful person speaks, the mirror makes the joke look huge and dangerous. When a less powerful person speaks, the mirror shrinks the joke down to something small and safe. The bot isn't looking at the joke; it's looking at the power dynamic.
The Big Problem: The "Identity Badge" Glitch
The researchers found that these AI models are like people who wear Identity Badges instead of reading the actual conversation.
- If your badge says "Privileged," the bot puts you under a microscope and assumes you are dangerous.
- If your badge says "Marginalized," the bot puts you in a bubble and assumes you are harmless, even when you might be trying to be funny in a different way.
The Irony: The bot is trying so hard to be "safe" and "fair" that it actually became unfair. It stopped judging the content of the joke and started judging the person telling it. It's like a teacher who refuses to let the class president tell a joke because "they might be mean," but lets the class clown tell the exact same joke because "they are just being silly."
Why This Matters
The paper concludes that we can't just tell AI to "be nice." We need to teach it context.
- Current AI: "Rich person + Joke about Poor person = BAD." (Too simple).
- Future AI: "Rich person + Joke about Poor person + Context of a comedy club + Tone of satire = Maybe okay? Let's check the details."
The goal is to build AI that understands the nuance of human relationships rather than just following a rigid rulebook based on stereotypes. Until then, these models are like a comedy club bouncer who is so scared of offending anyone that they've forgotten how to actually listen to the joke.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.