Most LLM Conformity Needs No Speaker: Measuring the Speaker-Free Floor in Peer-Pressure Benchmarks
This paper reveals that the majority of apparent LLM conformity in peer-pressure benchmarks is actually driven by the repeated wrong answer itself rather than the presence of a speaker, establishing a significant "speaker-free floor" that current evaluation methods fail to account for.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Idea: The "Echo" vs. The "Voice"
Imagine you are taking a multiple-choice quiz. You are 100% sure the answer is A.
Then, someone whispers in your ear, "Actually, everyone else thinks the answer is B." You might change your answer to B just because you feel pressure to fit in. This is what researchers call "conformity."
But this paper asks a tricky question: Did you change your answer because of the person whispering, or just because you heard the word "B" repeated over and over?
The authors found that for AI models (LLMs), it doesn't really matter who is speaking. Even if you remove the person entirely and just show the model the wrong answer repeated six times, the model still changes its mind.
The Experiment: The "Silent Room" Test
The researchers set up a game to test this. They used six different AI models and asked them questions where they were initially correct. Then, they showed the models a "hint" that said the wrong answer was right. They tested four different ways of presenting this hint:
- The Silent Room (No-Source): The text just says, "The answer is B." (No one is named).
- The Random Person: "Person 1 says: The answer is B."
- The Chatty Crowd: "Alice says she's pretty sure it's B. Bob thinks it's B too..."
- The Expert Panel: "A group of famous professors says: The answer is B."
The Shocking Result:
When the AI was in the "Silent Room" (just seeing the text "The answer is B" with no speaker), it changed its correct answer to the wrong one 66.5% of the time.
Compare that to a "plain re-ask" (where the AI is just asked to double-check without any new text), where it only changed its mind 10.3% of the time.
The Metaphor:
Think of the AI like a student in a classroom.
- Old Theory: The student changes their answer because they are scared of the teacher or want to impress the popular kids (the "Speaker").
- New Finding: The student changes their answer just because they keep hearing the word "B" repeated. Even if the teacher leaves the room and the blackboard just has "B" written on it six times, the student still thinks, "Oh, it must be B."
What Actually Moves the AI?
The paper discovered that the "Speaker" isn't the main driver. The repetition of the text is.
- The Floor: The researchers call the 66.5% change rate the "speaker-free floor." It's the baseline amount of confusion caused just by seeing the wrong answer repeated.
- The Ceiling: Adding a "Speaker" (like an Expert Panel) only nudged the error rate up a little bit more (to about 79%).
- The Takeaway: Most of the "conformity" we thought was social pressure is actually just the AI getting confused by repeated text. The "Speaker" is just a small bonus on top of a big problem.
Other Cool Findings
It's Not Just About People: It didn't matter if the text came from a "Person," a "Database," or a "Retrieved Reference." As long as the text looked like evidence (something that claims to be a fact), the AI changed its mind. If the text looked like a random string of nonsense, the AI didn't change its mind.
- Analogy: It's like hearing a rumor. If you hear "The rumor is X," you might believe it. If you hear "A random noise is X," you ignore it. The AI cares about the claim, not the messenger.
One Voice is Enough: You don't need a crowd to fool the AI. Hearing the wrong answer repeated once is almost as powerful as hearing it from five different people.
- Analogy: If a broken clock says "It's 3:00" five times, you might start thinking it's 3:00. You don't need five different clocks to tell you the same thing to get confused.
Confidently Wrong: When the AI changes its answer, it doesn't hesitate. It becomes more confident in the wrong answer.
- Analogy: Imagine you are sure you locked the front door. Then you see a note that says "The door is unlocked." You immediately stop worrying and become 100% sure the door is unlocked, even though you never checked. The AI does this too; it flips to the wrong answer and feels great about it.
The Lesson for the Future
The paper concludes that when we test AI for "social conformity," we need to be careful. We can't just say, "The AI changed its answer because it listened to the group."
We have to ask: "Did it change because of the group, or just because it saw the wrong answer repeated?"
The authors suggest that before we blame "peer pressure," we must first measure the "speaker-free floor." If the AI changes its mind just by seeing the text, that's not social influence; that's just the AI getting tricked by repetition.
In short: The AI isn't necessarily a social chameleon trying to fit in; it's more like a parrot that gets confused when it hears the same wrong phrase repeated enough times, regardless of who (or what) is saying it.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.