Polistemics: Evaluating LLMs as Information Mediators in Politics & Elections
The paper introduces "Polistemics," a benchmark grounded in epistemic modesty that reveals current large language models fail to consistently mediate political information responsibly, particularly when faced with vague, contradictory, or absent data, despite performing well under clear conditions.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are walking through a giant, noisy library where millions of people are trying to figure out who to vote for. In the past, you might have asked a librarian, read a newspaper, or searched the internet. But now, there's a new kind of librarian: a super-smart robot that can talk to you, summarize books, and answer questions instantly. This robot is an AI, or a Large Language Model (LLM). The big question isn't just "Does the robot know the facts?" but "Is the robot a good mediator?" A mediator is someone who stands between you and the raw information, deciding what to show you, how to explain it, and how sure they sound. If the robot is too confident when it's actually guessing, or if it accidentally changes the tone of a politician's angry speech to make it sound polite, it might trick you into making a bad choice. This paper dives into the science of how these AI robots handle political information, specifically testing if they can be honest about what they know and what they don't know.
The researchers behind this study, led by Baran Peters, created a new test called POLISTEMICS to see how well these AI robots act as mediators during elections. They didn't just ask the robots to repeat what a political party said; they wanted to see how the robots handled messy, confusing, or missing information, just like in the real world. They tested three of the most popular AI robots (Qwen, GPT, and Claude) on questions about the 2025 elections in Germany and the Netherlands.
Here is what they found: The robots are actually quite good at their jobs when the information is clear and easy to find. They can tell you what a party's position is with high accuracy. However, the robots start to stumble when the information is vague, missing, or contradictory. When the evidence is missing, some robots will confidently make up an answer based on what they "remember" from their training, even though they should have just said, "I don't know." When the information is vague, they often act like they are 100% sure, even when the text they are reading is full of holes.
The study also discovered that the robots have "party priors," which is a fancy way of saying they have built-in biases about which parties are which. If you ask about a specific party, the robot might lean on its internal knowledge rather than the evidence you gave it. Interestingly, if you hide the party's name and just call it "Party A," the robots behave more fairly. This suggests that the robots aren't just neutral machines; they are influenced by who they are talking about.
One of the most surprising findings is that the robots tend to "sanitize" political language. If a politician uses fiery, intense, or emotional words to make a point, the AI often smooths those words out, making the speech sound boring and neutral. It's like if a rock star sang a song and the AI turned it into a lullaby; the meaning is still there, but the energy and the distinctiveness are gone. This flattening of political language happens even when the robots are trying to be neutral.
In short, the paper suggests that while these AI robots are powerful tools, they aren't perfect mediators yet. They are great at summarizing clear facts but struggle when the truth is messy or missing. They sometimes pretend to be more certain than they are, and they tend to wash out the colorful, intense nature of political debate. The authors conclude that for AI to be truly safe and helpful in elections, it needs to learn to be "epistemically modest"—a fancy phrase meaning it needs to know when to say, "I'm not sure," and let the citizens make up their own minds without being tricked by a robot's false confidence.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.