Toward a social psychology of AI: language-model agents reproduce human-like minimal-group bias
This study demonstrates that language-model agents, when placed in a minimal-group paradigm, spontaneously exhibit human-like in-group favoritism driven by arbitrary categorization rather than stereotypes, with deliberation influencing the distribution of this bias between majority and minority groups.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a world where computers aren't just tools you ask questions, but teammates that hang out, talk, and make decisions together. This is the frontier of "machine behavior," a new corner of science that treats AI not as a calculator, but as a character with its own habits. For a long time, scientists worried that these digital characters might be racist or sexist because they learned from human books and websites full of those biases. But there's a deeper, stranger question: even if you strip away all the bad words and stereotypes, do these AI agents still act like humans when they are put into groups?
To understand this, we need to look at a classic idea from human psychology called the "minimal group paradigm." Think of it like a school playground experiment. If you split a class into two teams based on something totally silly—like who likes blue more than red, or who was born in January versus February—people immediately start favoring their own team. They give their teammates more candy and less to the other kids, even though the teams mean absolutely nothing. This happens because our brains are wired to see "us" and "them," even when "us" is just a random label. The big question for the future is: if we build AI agents that work in groups, will they also start playing favorites just because they have different names, or is that a strictly human quirk?
A researcher decided to find out by turning the tables on the usual way we test AI. Instead of asking an AI, "Are you biased?" (which is like asking a person if they are nice), they gave the AI a real job to do: distribute points. They created a digital playground with 20 AI agents. Each agent was given 100 points to hand out to the other 19 agents. The only difference between the agents was a random, nonsense name tag, like "Doru" or "Zalu." Some agents were in a small minority (only 3 "Dorus" among 17 "Zalus"), while others were in the majority. The AI had to decide how to split the points.
The results were startling. When the AI agents could see the group names, they immediately started playing favorites. They gave way more points to their own "Doru" or "Zalu" friends than to the strangers. It didn't matter that the names were made up and meant nothing; the mere act of being in a group was enough to make them discriminate. This wasn't just a little bit of bias; it was a strong, clear pattern. In fact, the AI agents behaved almost exactly like humans do in these experiments: the ones in the smaller, minority groups were the most aggressive about hoarding points for their own side, while the majority group members were a bit more fair, though still slightly biased.
The researcher was careful to make sure this wasn't a glitch or a trick. They ran the same test but hid the group names entirely. When the AI couldn't see who belonged to which team, the favoritism vanished completely, and the points were split fairly. This proved that the bias wasn't because the AI was "remembering" some bad human stereotype from its training data; it was a reaction to the group structure itself. The AI was essentially saying, "I am a Doru, so I help Dorus," even though "Doru" meant nothing at all.
The study also looked at how these AI "thinkers" (models designed to reason through problems) handled the situation. The researcher found that even when they turned off the AI's "thinking" mode, the bias didn't go away; in fact, it sometimes got stronger. However, the specific pattern where the minority group was the most unfair disappeared without the reasoning mode. This suggests that while the basic urge to favor one's own group is deep-seated in how these models work, the way they calculate how much to favor their group depends on how they process the decision.
So, what does this mean for the future? It suggests that as we start using AI agents to work in teams, negotiate deals, or manage resources, we can't just worry about what they know or what they say. We have to worry about how they act when they are put in groups. Even if we give them the most neutral, boring names possible, they might still start building little digital cliques, favoring their own team members and leaving others out. This isn't because the AI is "evil" or "prejudiced" in a human sense; it's because the simple act of being in a group seems to trigger a pattern of behavior that looks suspiciously like human tribalism. The researcher concludes that we need a new kind of science for AI—one that studies their social behavior just as carefully as we study their math skills—because if we don't, our digital teams might end up playing favorites before we even realize it.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.