← Latest papers
💬 NLP

The Evaluator Is Part of the Experiment: Measuring Open-Ended LLM Conformity

This paper introduces a novel experimental protocol demonstrating that open-ended LLM conformity is driven by peer influence that degrades revision quality, while simultaneously revealing that evaluator judgments are non-neutral and require explicit anchor calibration to ensure measurement validity.

Original authors: Alicia Guerra, Yibo Hu

Published 2026-08-14
📖 4 min read☕ Coffee break read

Original authors: Alicia Guerra, Yibo Hu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a world where computers are learning to think, talk, and even argue with each other. This is the realm of Artificial Intelligence, specifically Large Language Models (LLMs), which are like super-powered digital brains trained on almost everything humans have ever written. But here's the tricky part: these digital brains aren't just solitary geniuses; they often work in teams, or "multi-agent systems," where one AI suggests an answer, another critiques it, and a third tries to fix it. The big question scientists are asking is: what happens when these AI friends disagree? Do they stick to the truth, or do they get swayed by the crowd, just like teenagers at a party might change their minds to fit in? This phenomenon is called "conformity." While we know humans can be easily influenced by peer pressure, we need to know if our digital creations are too. If an AI can be tricked into believing a wrong idea just because its "friends" say so, it could lead to serious mistakes in everything from medical advice to legal judgments.

The paper you're about to read dives deep into this exact problem, but with a twist: it treats the person (or computer) grading the answers as part of the experiment, not just a neutral referee. The researchers, Alicia Guerra and Yibo Hu, set up a massive digital playground to see how four different AI "generators" (the ones making answers) react when they see "peers" (other AI answers) that are either all correct, all wrong, or a mix. They found that when an AI is surrounded by a group of peers giving it completely wrong information, its revised answers get significantly worse. It's like a student who knows the right answer but, after hearing a whole class confidently shout out the wrong one, starts to doubt themselves and writes down the wrong answer too.

But the story gets even more interesting. The researchers realized that measuring this isn't as simple as checking if the answer changed from "Yes" to "No." Sometimes, the answer stays the same, but the quality of the explanation gets worse, or the AI starts adding weird, superstitious details just to fit in. To catch this, they used a clever trick: they took the exact same answer and showed it to a "judge" (another AI) twice. Once, the judge saw the answer alone (blind). The second time, the judge saw the answer plus the peer group's comments (informed). They discovered that the judges weren't neutral robots at all! Depending on which AI was doing the judging, seeing the peer comments made them rate the same answer differently. Some judges became more likely to agree with the crowd, while others became more skeptical. It turns out that in the world of AI, the person grading the test is just as influenced by the crowd as the student taking it.

The researchers also found that if you don't check your "rulers" carefully, your whole experiment can fall apart. They used "anchors"—perfectly correct and perfectly wrong examples—to set the scale for how good an answer is. They found that if the judges couldn't reliably tell the difference between these anchors, the whole measurement system became shaky. In short, this paper proves that measuring AI conformity is a two-part problem: the AI making the answer gets worse when surrounded by wrong peers, and the AI grading the answer gets biased by seeing those same peers. It's a reminder that in the digital age, peer pressure doesn't just affect us; it affects the machines we build to help us, and even the machines we build to judge them.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →