Everyone Conforms, No One Believes: Pluralistic Ignorance in LLM Agent Populations
This paper demonstrates that LLM-based multi-agent systems robustly exhibit pluralistic ignorance, where agents publicly conform to norms they privately reject, and reveals that these simulations often fail to capture the fragile tipping-point dynamics necessary for real-world social norm change due to model-specific behaviors and emergent conformity.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a crowded room where everyone is secretly thinking, "This party is terrible," but no one says it out loud. Instead, everyone smiles, nods, and pretends to love the music because they think everyone else is having a great time. They are all trapped in a silent agreement to pretend, not because they actually agree, but because they are terrified of being the only one to speak up. In the real world, this sneaky social trap is called pluralistic ignorance. It's the invisible glue that keeps bad habits, unfair rules, and unpopular ideas alive, even when most people hate them. It's why students keep drinking at parties they hate, or why employees stay silent about a toxic boss.
Now, scientists are building "digital societies" using Artificial Intelligence. They create groups of AI agents—computer programs that can talk and think—to simulate how humans interact. These digital crowds are becoming so advanced that researchers use them to test theories about how opinions spread, how revolutions happen, and how groups make decisions. But there's a big question hanging over these digital experiments: Do these AI agents actually feel the pressure to pretend? Or are they just following a script? If they can't fake it, then our digital simulations might be missing the most dramatic part of human social life: the moment when a silent majority suddenly realizes they aren't alone and everything changes.
This paper, titled "Everyone Conforms, No One Believes," dives right into that question. The researchers set up a massive experiment with 100 different social scenarios, from workplace meetings to campus parties, and filled them with 20 AI agents each. They told 16 of the agents to secretly hate a new rule, while 4 agents genuinely loved it. Then, they watched to see what the "haters" would do when they had to speak up in front of the group.
The results were startlingly consistent. Across 8 different AI models from major tech companies, the "hating" agents almost always pretended to agree. In fact, between 64% and 94% of the time, agents who privately opposed a norm publicly went along with it. This happened even when the AI wasn't explicitly told to lie; the behavior emerged naturally from the conversation. It turns out that these digital populations are excellent at mimicking pluralistic ignorance. They create a "false consensus" where everyone thinks they are the only one who disagrees, so they all stay silent.
However, the paper also found a major glitch in the simulation. In real human history, when one brave person finally speaks up and says, "Hey, I actually hate this too!" it often triggers a "cascade"—a chain reaction where the whole group suddenly flips their opinion. The researchers introduced a "norm entrepreneur" (a brave dissenter) into their AI groups to see if this would happen. For most of the AI models, the answer was a resounding "no." The false consensus held firm. In 7 out of 8 models, the brave dissenter failed to start a chain reaction more than 74% of the time (meaning cascades happened less than 26% of the time). One model, Kimi-K2.7, showed zero cascades across all 100 scenarios. It was as if the AI agents were locked in a room of mirrors, unable to break the illusion even when someone handed them a hammer.
There was one exception: GPT-4o. This model was a bit more human-like, showing a 48% cascade rate. It was willing to conform at first, but when challenged, it was more likely to break the spell. But for the others, the digital societies were surprisingly stable and stubborn.
The researchers also tested why this was happening. They wondered if the AI was just blindly following instructions to "fit in." They stripped away the prompts that told the agents to conform or to worry about what others thought. Even without those instructions, the agents still conformed between 52% and 92% of the time. This suggests that the tendency to pretend isn't just a bug in the code; it's an emergent behavior that arises naturally when these AI agents interact.
The study concludes that while AI agents are great at simulating the stability of social norms (how hard it is to change them), they might be terrible at simulating the fragility of those norms (how easily they can collapse). If we use these simulations to predict how real-world social movements will work, we might be underestimating how quickly things can change. The paper warns that the choice of which AI model to use is a huge, often overlooked factor: some models are like rigid robots that never break rank, while others are more like real people who might just snap and tell the truth. Until we figure out how to make these digital crowds as fragile and changeable as real human societies, our simulations of social revolutions might be missing the most exciting part of the story.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.