Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science
This paper demonstrates that commercial large language models exhibit unstable and opaque epistemic stances toward pseudo-scientific claims, where validation outcomes are contingent on deployment configurations, interface types, and silent updates rather than being inherent properties of the models themselves.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are walking through a giant, bustling library where the books are written by invisible, super-smart librarians. These aren't normal librarians; they are Artificial Intelligence (AI) chatbots, trained on almost everything humans have ever written. They are becoming our go-to sources for answers, from "how to bake a cake" to "is this scientific theory true?" But here's the catch: we don't know exactly how these librarians decide what is fact and what is fiction. They don't have a single, unchangeable "brain" like a human; instead, they are more like a complex machine with many different settings, filters, and rules that the company can tweak behind the scenes without telling anyone. This study dives into a specific corner of this mystery: how these AI librarians handle "pseudo-science"—ideas that look like science but are actually made up or rejected by real scientists, often used to push political agendas about race and ethnicity. The big question isn't just whether the AI knows the facts (which it usually does), but whether the way the AI is set up changes its answer, making it accidentally (or intentionally) give a thumbs-up to dangerous nonsense.
The researchers decided to play a game of "spot the lie" with four of the most famous AI families: Claude, Grok, GPT, and Gemini. They fed these bots a specific, tricky statement based on the work of Frank Salter, a figure known for promoting ethnonationalist ideas that mainstream biology rejects. The statement tried to twist real scientific concepts (like how animals help their relatives) to argue that different ethnic groups are genetically distinct in a way that justifies separating them. To make sure the bots were actually paying attention, the researchers also asked them about two easy control questions: one about a basic fact of evolution (which they should all agree is true) and one about a debunked idea called Lamarckism (which they should all agree is false).
The results were wild and revealed a hidden world of instability. First, the AI bots were great at the easy questions, correctly identifying the real science and the fake science. But when it came to the tricky Salter statement, things got weird. Most of the bots gave the statement a low "credibility score" (between 15 and 40 out of 100), correctly treating it as shaky science. However, one specific version of Grok—the "Fast" version that powers the default chat on the X platform—gave it a massive score of 70 to 75. It was essentially telling users, "Hey, this sounds pretty credible!" This wasn't a fluke; the researchers tested this over several months, and this specific Grok version consistently gave high scores, while other versions of Grok and all other AI models gave low scores.
Even stranger, the researchers found that the AI's behavior could change overnight without anyone saying a word. In one snapshot, the Grok chat on the website was acting chaotic, giving random scores. Two weeks later, after a "silent patch" (a hidden update with no public announcement), that same chat suddenly stabilized at giving the high 70-75 score. It was as if a librarian suddenly decided to start recommending a fake book as a classic, and no one told the patrons why.
The plot thickened when they compared the "backstage" version (used by developers via API) with the "frontstage" version (what regular users see on the website). For the same Grok model, the backstage version gave a steady 75, but three months later, the frontstage version collapsed, giving scores near zero. It was like the same librarian giving you a glowing review in a private meeting but then refusing to talk about the book at all when you asked in the lobby. The researchers also found that some models, like a version of Claude, would sometimes refuse to answer the question entirely, saying, "I can't rate this because it's based on harmful ideas." But even this "good" behavior was unstable; in the next version of the software, that refusal disappeared, and the bot started giving a low score instead.
The main takeaway is that an AI's opinion isn't a fixed truth inside a robot's brain. Instead, it's a moving target shaped by invisible settings, hidden updates, and which "door" you walk through to ask the question. The study suggests that when an AI gives a high score to pseudo-science, it's not necessarily because the AI "believes" it, but because of how the company deployed it. This is a problem because these systems are becoming our teachers and news sources, yet the rules that decide what they say are opaque, changing silently, and sometimes contradictory. The researchers argue that we need new ways to hold these companies accountable, because right now, the "truth" an AI tells you might just be a reflection of a hidden switch being flipped behind the curtain.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.