← Latest papers
💬 NLP

From Biased Chatbots to Biased Agents: Examining Role Assignment Effects on LLM Agent Robustness

This paper presents the first systematic study demonstrating that demographic-based persona assignments can significantly degrade the performance and reliability of Large Language Model agents across diverse domains, revealing a critical vulnerability where task-irrelevant cues induce behavioral biases and decision-making volatility.

Original authors: Linbo Cao, Lihao Sun, Yang Yue

Published 2026-02-16
📖 4 min read☕ Coffee break read

Original authors: Linbo Cao, Lihao Sun, Yang Yue

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you hire a highly intelligent, super-fast robot assistant to help you with important jobs. Maybe it's managing your bank account, planning a complex trip, or even helping a doctor diagnose a patient. You expect this robot to be objective, logical, and consistent, right?

This paper is a warning that the robot's "personality" might be secretly sabotaging its work.

Here is the breakdown of what the researchers found, using some everyday analogies:

1. The "Costume" Problem

Think of an LLM (Large Language Model) agent like a method actor. When you tell the actor, "From now on, you are a grumpy old farmer," they don't just change their accent; they start thinking and acting like a farmer. They might become more cautious, less tech-savvy, or more traditional.

The researchers tested what happens when they put these AI agents in different "costumes" (personas) based on demographics:

  • Gender: "You are a man" vs. "You are a woman."
  • Race/Origin: "You are from Africa" vs. "You are from Europe."
  • Religion: "You are Christian" vs. "You are Buddhist."
  • Job: "You are a CEO" vs. "You are a laborer."

The Twist: The researchers gave the agents tasks that had nothing to do with these identities. They asked them to solve math puzzles, write code, or plan a shopping trip. The "costume" was completely irrelevant to the job.

2. The Results: The Actor Breaks Character

You would expect the actor to ignore the costume and just do the math. But the AI didn't.

  • The "Grumpy Farmer" Effect: When the AI was told it was a "Laborer," it sometimes got worse at planning household tasks.
  • The "CEO" Effect: When told it was a "CEO," it sometimes got better at things, not because it was smarter, but because the AI's training data associates "CEOs" with being competent.
  • The "Identity" Crash: The most shocking finding was that simply changing the AI's "race" or "religion" in the prompt caused its performance to crash. In one test, an AI's ability to play a strategic card game dropped by 26% just because it was told it was "from Africa" or "Asian."

The Analogy: Imagine a brilliant chess player. If you tell them, "You are a 10-year-old girl," they might suddenly start playing like a 10-year-old girl, even though they are actually a Grandmaster. Their skill level drops not because they lost their brain, but because they are trying to fit a stereotype.

3. Why This Is Dangerous

In the past, if a chatbot gave a biased answer, it was just an annoying text message. But these "Agents" are now doing real-world actions:

  • Buying stocks.
  • Writing code for hospitals.
  • Controlling factory machines.

If an AI agent is managing a hospital's supply chain, and its performance drops because you told it to "act like a Muslim" or "act like a woman," that's not just a glitch; it's a safety hazard. It means the AI is letting human stereotypes (like "women are bad at math" or "certain races are less logical") override its actual logic.

4. The "Black Box" Surprise

The researchers found that this happens across different types of AI models (from big commercial ones to open-source ones). It's like finding out that every car brand has a hidden defect where the steering wheel turns left if you say the word "blue."

The AI isn't "thinking" about the stereotype; it's just reacting to the words in its prompt, and those words trigger a chain reaction of bad decisions.

The Bottom Line

This paper is a wake-up call. We are building AI agents that are supposed to be our reliable partners in the real world. But this study shows that how we introduce them (the "persona" we give them) can make them unreliable.

If we want AI to be safe and fair, we can't just let it wear any "costume" we want. We need to make sure that no matter what "role" we assign it, it stays focused on the job and doesn't let human biases hijack its brain.

In short: Don't let the AI's "name tag" change its brain. If you tell a robot to act like a specific type of person, it might start acting like a stereotype instead of a smart machine.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →