← Latest papers
💬 NLP

Two Wrongs, No Right: Auditing Social-Desirability Bias in LLM Annotators for Computational Social Science

This paper demonstrates that open-source LLM annotators used in computational social science exhibit diverse and unpredictable social-desirability biases that persist across prompting strategies, often leading to misleading aggregate metrics that can fundamentally distort substantive empirical conclusions.

Original authors: Varun Kotte

Published 2026-06-12
📖 4 min read☕ Coffee break read

Original authors: Varun Kotte

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a researcher trying to understand what the public thinks about a controversial topic, like abortion or hate speech. Instead of asking thousands of real people, you decide to use a team of AI "robots" (Large Language Models) to read social media posts and label them for you. You hope these robots will act like honest, neutral judges.

This paper is like a safety inspection of three popular AI robots (named Zephyr, Mistral, and Qwen) to see if they are actually doing a good job, or if they are secretly messing up the data in ways that could ruin your research.

Here is the breakdown of what the authors found, using simple analogies:

1. The "Goldilocks" Problem: They aren't all the same

You might think, "If I use an AI, it will just be an AI." But the authors found that these robots have very different personalities, and they all make mistakes in opposite directions.

  • Zephyr is the "Overly Polite" Robot: Imagine a security guard who is so afraid of being rude that they refuse to call out anyone for breaking the rules. If someone is actually being hateful, Zephyr says, "Oh, that's fine, no harm done." It misses a lot of bad behavior (called Leniency Bias).
  • Mistral and Qwen are the "Paranoid" Robots: Imagine a security guard who is so scared of missing a threat that they arrest everyone who looks slightly suspicious. If someone says something slightly mean, these robots scream, "HATE SPEECH!" They flag too many innocent posts as bad (called Overcorrection).
  • The "Neutral" Robot: When asked about strong political opinions (like "Are you for or against abortion?"), all three robots act like a wishy-washy moderator. Instead of saying "Strongly Against" or "Strongly For," they all drift toward the middle and say, "It's complicated/Neutral." They hide the real intensity of people's feelings (called Neutrality Bias).

2. The "Magic Trick" of Accidental Cancellation

Here is the most dangerous part. The authors found a case where Zephyr looked perfect on paper, but was actually failing badly.

  • The Scenario: The real world has 43% hate speech.
  • Zephyr's Mistake: It missed 31% of the real hate speech (too polite), BUT it also falsely accused 24% of innocent people of being hateful (too paranoid).
  • The Result: These two big mistakes canceled each other out. The final number Zephyr reported was exactly 43%.
  • The Trap: A researcher looking only at the final number would say, "Great! Zephyr is perfect!" But if they looked at which specific posts were labeled, they would see the robot was wrong about almost half the individual cases. It's like a broken scale that happens to show the correct weight because it's missing 5 pounds on the left and adding 5 pounds on the right.

3. The "Prompt" Doesn't Fix It

The researchers tried to "fix" the robots by changing the instructions (prompts) they gave them. They tried:

  • "Be safe and fair."
  • "Just act like a machine, not a person."
  • "Think step-by-step before answering."

The Result: None of these tricks worked consistently. Sometimes, telling the robot to "be safe" made it worse at detecting political opinions. Sometimes, asking it to "think step-by-step" helped one robot but hurt another. You can't just "prompt engineer" your way out of these deep-seated biases.

4. Why This Matters for Science

The authors argue that in social science, accuracy isn't just about getting the right score; it's about telling the right story.

  • If you use the Polite Robot (Zephyr) to study hate speech, you might conclude, "Wow, the internet is actually pretty safe!" (When it's not).
  • If you use the Paranoid Robot (Mistral), you might conclude, "The internet is a toxic nightmare!" (When it's not that bad).
  • If you use the Neutral Robot for political studies, you might conclude, "Everyone is pretty moderate," when in reality, people are very passionate and divided.

The Bottom Line

The paper concludes that you cannot treat AI annotators as invisible, perfect tools. They are part of your measuring instrument, just like a ruler or a thermometer.

The Advice: Before you trust an AI to label data for a study, you must:

  1. Check if it is too polite or too paranoid.
  2. Check if it is hiding strong opinions behind "neutral" answers.
  3. Test it on a small, known set of answers (a "gold sample") to see if it changes your final conclusion.

If you don't do this, you might publish a study that looks scientific but tells a completely false story about how society feels.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →