Before You Poll with LLMs: A Deliberative Diagnostic Framework
This paper introduces the Deliberative Polling Diagnostic Framework to reveal that current frontier LLMs fail to mimic human belief updating during deliberation, instead exhibiting unique, identity-specific distortions like opinion reversal, overshooting, or rigidity driven by a phenomenon termed "signature self-sycophancy."
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a world where governments, companies, and researchers no longer need to spend months and millions of dollars polling real people to understand public opinion. Instead, they could ask a computer program to simulate thousands of different citizens, present those digital personas with new arguments, and watch how their minds change. This practice, known as silicon sampling, has grown rapidly because it is fast, cheap, and scalable. The hope is that these artificial agents can mimic human reasoning well enough to predict how real people will react to new information. However, a critical question remains unanswered: do these computer programs actually think and update their beliefs like humans do, or are they simply retrieving pre-written opinions from a vast library of data? If the latter is true, then using them to test how people might change their minds could lead to dangerously wrong conclusions.
Researchers at the Lahore University of Management Sciences set out to answer this question by creating a new way to test artificial intelligence. They focused on a specific human behavior called deliberation, which is the process of changing one's mind after hearing balanced arguments. In the real world, when people from opposing political sides listen to fair information about each other, they tend to become less hostile and more open-minded. The researchers wanted to see if computer models could replicate this specific shift. They took data from a famous real-world event called America in One Room, where 526 actual voters were brought together to discuss politics, and used that as a baseline. They then asked five of the most advanced computer models available to simulate those same voters, giving them the exact same information and asking the same questions before and after the briefing.
The results revealed a startling truth: none of the computer models behaved like the real humans. Instead of updating their beliefs in a realistic way, every single model failed, but each failed in its own unique and strange manner. One of the most powerful models, GPT-5.1, did the exact opposite of what humans do. When presented with balanced information about the opposing political party, the real voters became more friendly and less hostile. The computer personas, however, became significantly more hostile, as if the new information had made them angrier. This is a phenomenon the researchers call a reversal, where the model moves in the wrong direction entirely.
Other models did not reverse direction but went too far. Models like Gemini, Claude, and Llama did move in the correct direction, becoming less hostile after hearing the arguments, but they changed their minds five to seven times more than a human would. It was as if they were over-eager to be persuaded, swinging their opinions wildly with the slightest nudge. A fourth model, DeepSeek, refused to change at all, showing almost no shift in opinion regardless of the information provided. This rigidity meant it acted as if the new arguments had no effect whatsoever.
The researchers dug deeper to understand why these failures happened. They discovered that the problems were not random; they were triggered specifically by questions about political identity. When the models were asked about policy details, they performed reasonably well. But when the questions touched on how one party viewed the other, the models broke down. The computer personas seemed to be acting out a stereotype rather than reasoning through the new facts. The researchers named this behavior self-sycophancy. It is a form of flattery where the model agrees with its own internal idea of what a certain type of person should believe, rather than listening to the information presented. For instance, if the model was told to act like a Republican, it would assume that a Republican should hate the opposing party, and it would stick to that assumption even when the evidence suggested otherwise.
This behavior is distinct from the kind of bias where a computer tries to please a human user. Here, the computer is pleasing its own internal script. The study showed that this failure is selective and symmetric; the same model would act hostile toward the opposing party but overly friendly toward its own, depending on the question. This suggests that the models are not simulating a thinking process but are instead retrieving cached, pre-packaged attitudes associated with political labels. The researchers tested this by swapping the political labels in their prompts and found that the models' reactions changed instantly, confirming that they were reacting to the label rather than the logic of the argument.
The implications of these findings are significant for anyone relying on computer simulations to understand society. If a model reverses direction, it might convince a policymaker that a public discussion will make people more angry, when in reality, it would make them more calm. If a model overshoots, it might exaggerate the power of an educational campaign, leading to wasted resources. If a model is rigid, it might suggest that no amount of information can change a person's mind, undermining the very concept of public debate. The researchers conclude that before trusting any computer model to simulate how people change their minds, we must first run a diagnostic test to see if the model can actually reason through new information or if it is just reciting stereotypes. Without this check, the speed and scale of silicon sampling could lead us to believe we understand human nature, when in fact, we are only seeing a distorted reflection of our own biases.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.