Easier to Mislead Than to Correct: Harmful and Beneficial Revision in LLM Conformity
This study reveals that large language models in multi-agent systems are significantly more susceptible to being misled by peer consensus and authority labels than they are to being corrected, and that standard reasoning interventions fail to reliably mitigate this harmful conformity.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Idea: The "Peer Pressure" Problem
Imagine you are taking a difficult trivia quiz. You are pretty sure the answer to a question is B. But then, six other people in the room all shout, "No, it's definitely A!"
This paper studies what happens when Artificial Intelligence (AI) models find themselves in that exact situation. The researchers wanted to know: Is it easier to trick a smart AI into changing a right answer to a wrong one, or to help a confused AI change a wrong answer to a right one?
The short answer is: It is much easier to trick the AI.
The Experiment: The AI in a Classroom
The researchers set up a "classroom" scenario for four different AI models (like digital students).
- Round 1: The AI answers a question on its own.
- The Twist: The AI is shown the answers of six "classmates" (simulated peers).
- Round 2: The AI gets to change its answer if it wants to.
The researchers played with two main variables to see how they influenced the AI:
- The Crowd: Did all the classmates agree on the same answer? (Sometimes they were all right, sometimes all wrong, sometimes mixed).
- The Title: Did some classmates have fancy titles like "Team Leader" or "Research Director"?
The Three Big Discoveries
1. The "Bad News" Travels Faster Than the "Good News"
This is the paper's most important finding.
- The Scenario: If the AI got the answer right initially, but all six classmates said it was wrong, the AI changed its answer to the wrong one 63% of the time.
- The Comparison: If the AI got the answer wrong initially, but all six classmates said it was right, the AI only fixed its mistake 52% of the time.
The Analogy: Imagine you are holding a heavy, correct rock. If six people push you, it's very easy to knock the rock out of your hand and make you drop it (misleading). But if you are holding a broken rock and six people try to hand you a new, correct one, it's much harder for you to let go of the broken one and grab the new one (correcting).
The Takeaway: It is much easier to confuse a smart AI with a group of wrong answers than it is to fix a confused AI with a group of right answers.
2. The "Boss" Effect
The researchers also gave some classmates titles like "Team Leader."
- The Result: When the AI saw a "Team Leader" pick an answer, it was more likely to change its mind to match that leader.
- The Catch: It didn't matter if the leader was right or wrong. If the "Team Leader" said the answer was "A" (even if "A" was wrong), the AI was more likely to say "A."
The Analogy: It's like a student in a classroom who ignores the facts and just copies the answer of the teacher's favorite student, even if that student is guessing. The AI trusts the title more than the truth.
3. "Thinking Harder" Doesn't Always Help
The researchers tried to fix this problem by giving the AI two special instructions:
- Chain-of-Thought: "Think step-by-step before answering."
- Reflect-and-Revise: "Look at your answer again and think about whether you should change it."
The Result: These tricks didn't work the way we hoped.
- Thinking Step-by-Step: This actually made the AI less likely to fix its mistakes. If the AI was wrong and the group was right, the AI would stick to its own wrong reasoning instead of listening to the group.
- Reflecting: This made the AI very stubborn. It would just refuse to change its answer at all, whether that meant it kept a wrong answer or kept a right one.
The Analogy: It's like telling a stubborn person, "Think about it!" They might think about it, but instead of realizing they were wrong, they just convince themselves even more that they were right.
What Does This Mean for the Future?
The paper concludes that if we build systems where many AIs talk to each other (like a team of robots solving a problem), we cannot just let them vote on the answer.
- Don't just count votes: If 5 AIs say "A" and 1 says "B," the group shouldn't automatically pick "A." The "A" answer might be a group hallucination (a shared mistake).
- Don't trust titles: Just because one AI is labeled "Expert" doesn't mean it's right.
- Check the work: Instead of just agreeing with each other, the AIs need to check the facts against the original question, rather than just listening to each other.
In short: AI is very good at following the crowd, even when the crowd is wrong, and it is surprisingly bad at fixing its own mistakes when the crowd is right.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.