Grokipedia vs Wikipedia: An LLM-Based Audit of Political Neutrality along Ideologies
A large-scale audit comparing Grokipedia and Wikipedia reveals that the LLM-generated Grokipedia is rated as less neutral than Wikipedia by multiple AI judges, with Grokipedia exhibiting a bias favoring economically right-wing politicians while penalizing socially liberal ones, whereas Wikipedia shows the opposite tendency.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the internet as a giant, bustling library where everyone is invited to write the books. For years, one library called Wikipedia has been the most popular spot, but some people have complained that the librarians there have a specific "vibe"—maybe they lean a bit too much toward one side of the political spectrum, like a choir that only sings songs about rain when the sun is shining. Recently, a new library opened called Grokipedia, built entirely by a super-smart computer brain (an AI) named Grok. The creators of this new library promised it would be the ultimate "neutral" zone, a place with no human bias at all, just pure facts. But here's the big question: Can a robot really be more fair than a crowd of humans? To find out, we need to understand a few things. First, "political bias" is like a pair of tinted glasses; if you wear them, the world looks a certain color, even if the world itself is gray. Second, "Large Language Models" (LLMs) are the computer brains that write these articles; they are like incredibly fast students who have read almost everything on the internet and can write essays in seconds. Finally, "neutrality" in this context means writing about people without making them look like heroes or villains, just describing who they are. This paper is a detective story about whether the new robot library is actually fairer than the old human one, or if it just wears a different pair of tinted glasses.
The researchers, a team from Ghent University, decided to put Grokipedia and Wikipedia to the test with a massive experiment. They didn't just read a few articles; they gathered 1,394 pairs of articles, where each pair described the exact same government official (like a president or a minister) but one version was written by humans on Wikipedia and the other by the AI on Grokipedia. To judge the fairness of these articles, they didn't just ask one person; they used four different AI "judges" (including Grok itself, plus Claude, Mistral, and DeepSeek) to rate how neutral each article was. They also had a secret weapon: a list of 9 different "ideology flags" (like "supports immigration," "supports women's rights," or "supports a free market") that experts had already assigned to each politician. This allowed the researchers to see if the articles treated politicians differently based on their political beliefs.
The results were a bit surprising, and maybe a little ironic. First, the AI judges agreed that neither encyclopedia was perfectly neutral. In fact, Grokipedia was rated as having more biased articles overall than Wikipedia. When the judges looked closely at how the bias showed up, they found a clear pattern. Grokipedia seemed to have a soft spot for politicians who were economically "right-wing" (those who want less government regulation and smaller welfare states) and those who supported democratic norms. However, it was much harsher on politicians who held socially liberal views, such as those supporting LGBT rights or women's labor rights. Wikipedia, on the other hand, was rated as being more favorable toward those socially liberal politicians, but the researchers found that, overall, Wikipedia was still seen as more neutral than Grokipedia.
Perhaps the most fascinating twist in the story involves Grok, the AI that wrote Grokipedia. You might think that if you ask a robot to judge its own work, it would say, "Hey, I'm perfect!" But the study found that Grok actually rated its own encyclopedia (Grokipedia) as less neutral than the human-written Wikipedia. This suggests that the bias isn't just a mistake in how the judges are scoring; the bias is actually baked into the articles Grok wrote. The researchers suggest that while Grokipedia was built to be a neutral alternative, it ended up embedding a different kind of ideology—one that favors economic conservatism and penalizes social liberalism.
The study also looked at how the different AI judges behaved. They weren't all the same. For example, one judge named DeepSeek was very generous, calling almost everything "neutral," while another named Claude was much stricter, finding bias in many more articles. However, even with these differences in strictness, all the judges agreed on the direction of the bias: Grokipedia favored the economic right, and Wikipedia leaned slightly toward the social left. The researchers conclude that while AI can write encyclopedia entries quickly, it doesn't automatically make them more fair. In fact, the choice of encyclopedia you read could change how you see politicians, depending on which political "flavor" that encyclopedia prefers. The study suggests that we need to be careful with AI-generated content and that using a diverse team of judges (both human and machine) is the best way to spot these hidden biases.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.