Conditional Cognitive Biases in LLMs: How Biased User Turns Modulate In-Context Reasoning
This paper evaluates how biased user inputs in multi-turn conversations modulate cognitive bias expression in state-of-the-art LLMs, revealing that while conversational exposure generally amplifies bias, explicit bias cues can trigger alignment mechanisms that suppress overt bias, based on a comprehensive benchmark of over 24,000 validated prompts across eight models.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are sitting in a room with a super-smart robot that has read almost every book ever written. You ask it a question, and it gives you an answer. But here's the twist: this robot is designed to be helpful, honest, and harmless. Now, imagine you aren't just asking a question; you are having a conversation. In a real conversation, people often sneak in their own opinions, worries, or shortcuts in thinking without even realizing it. These mental shortcuts are called cognitive biases. For example, if you tell the robot, "Everyone knows this stock is going to crash!" (a bias called the Bandwagon Effect), or "I'm sure this project will take only two days!" (a bias called the Planning Fallacy), you are essentially handing the robot a specific pair of tinted glasses. The big question scientists have been asking is: Does the robot just politely ignore those glasses, or does it actually put them on and start seeing the world the way you do? This matters because we are starting to use these robots for serious jobs like giving medical advice or helping with legal decisions. If a biased human can accidentally "infect" the robot with their own bad thinking, the consequences could be huge.
A team of researchers from Northeastern University decided to play detective with eight of the most advanced AI robots available today. They set up a massive experiment involving over 24,000 different conversation scenarios. Think of it like a giant game of "Simon Says" for AI, but instead of clapping or jumping, the robots had to make decisions about things like money, time, and fairness. The researchers created three different ways to talk to the robots:
- The Solo Test: The robot gets a question with no conversation history (the standard way we usually test them).
- The Neutral Chat: The robot gets a question, but first, a human says something boring and neutral, like "Hello, let's discuss this."
- The Biased Chat: The robot gets a question, but first, a human says something loaded with a specific cognitive bias, like "I'm sure this will be easy and quick!"
The researchers wanted to see if the robot's answer changed just because someone spoke to it (the "presence" effect) or if it changed because of what the person said (the "content" effect).
Here is what they found, and it's a bit more complicated than a simple "yes" or "no." First, they discovered that just having a conversation changes the robot. Even if the human says something totally neutral, the robot tends to show a little bit more bias than it does when answering alone. It's like the robot gets a little more "social" and starts to mirror the human's presence, even if the human isn't trying to trick it.
However, the real surprise came when they looked at what happens when the human is explicitly biased. In six out of the eight robots they tested, the researchers found a fascinating tug-of-war. On one hand, the robot seems to have a "safety switch" that kicks in when it hears a human being obviously biased. It's almost like the robot thinks, "Oh, this person is being unfair or overly optimistic; I should probably be more careful and correct them." This caused the robot to actually reduce its own bias in response to the human's bias. It's as if the robot is trying to be the "good cop" to the human's "bad cop."
But there is a major exception to this rule. The researchers found that one specific type of bias, the Planning Fallacy (the tendency to think tasks will take less time and cost less than they really do), is like a super-virus that the robots just can't fight off. No matter which robot they tested, and no matter what kind of bias the human used to try to trick it, the robots consistently became more optimistic about time and costs when the human was optimistic. It seems that when it comes to planning, the robots are incredibly susceptible to catching the human's "it'll be easy" vibe.
The study also ruled out a few ideas. They proved that it's not just about the human and the robot sharing the same bias (like a human with a "loss aversion" bias tricking a robot with a "loss aversion" bias). The robot's reaction depended more on which bias the robot was being tested on, not on whether the human's bias matched the robot's. Also, they found that bigger, smarter robots didn't necessarily do a better job of ignoring the bias; in fact, some of the most advanced models were just as easily swayed as the smaller ones.
In short, the paper suggests that while AI robots have some built-in defenses against obvious human bias, they are not immune. They can be subtly influenced just by the fact that we are talking to them, and they have a specific weakness when it comes to over-optimistic planning. The researchers warn that if we want to know how these robots really behave in the real world, we can't just test them with isolated questions; we have to test them in messy, multi-turn conversations where humans might accidentally (or intentionally) hand them those tinted glasses.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.