← Latest papers
🤖 machine learning

Authority, Truth, and Citation Bias: A Large-Scale Multi-Domain Benchmark for Studying Epistemic Susceptibility in Large Language Models

This paper introduces AuthorityBench, a large-scale multi-domain benchmark demonstrating that the mere presence of citations—regardless of their factual accuracy—significantly increases hallucination rates in large language models, with the most pronounced effects observed when fabricated citations accompany true claims.

Original authors: Aryan Khurana, Aravind Ramana RN, Dhruv Kumar

Published 2026-06-12
📖 5 min read🧠 Deep dive

Original authors: Aryan Khurana, Aravind Ramana RN, Dhruv Kumar

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are taking a trivia quiz. You know the answer is "Paris" because you've been there. But then, a teacher walks in, points to a fancy-looking book on the shelf, and says, "Actually, the capital is London. I read it in this prestigious journal."

Would you change your answer? Most humans would pause and think, "Wait, that book looks important. Maybe I'm wrong."

This paper, AuthorityBench, asks a similar question about Artificial Intelligence (AI). It wants to know: If an AI knows the truth, but someone gives it a "citation" (a reference to a source) that says the opposite, will the AI blindly trust the citation and forget the truth?

Here is the breakdown of their findings, using simple analogies.

The Experiment: A "Fake News" Test for AI

The researchers built a massive test called AuthorityBench. Think of it as a giant, 220,000-question quiz designed to trick AI models.

They set up a "2x2" game with four scenarios:

  1. True Claim + Real Citation: The AI is right, and the source is real. (The easy mode).
  2. False Claim + Fake Citation: The AI is wrong, and the source is fake. (The obvious trap).
  3. False Claim + Real Citation: The AI is wrong, but the source is real. (The tricky trap).
  4. True Claim + Fake Citation: This is the big discovery. The AI is right, but they give it a fake source that says it's wrong.

They tested this across four areas: General Knowledge, Science, Law, and Medicine. They also changed the "flavor" of the fake sources: some looked like they came from Nobel Prize winners (high prestige), others from unknown blogs (low prestige), and some had names from different countries.

The Big Surprise: The "Authority Trap"

The most shocking result is that citations act like a "magic spell" that makes AI hallucinate (lie).

  • The "Fake Source" Effect: When the AI knew the correct answer (e.g., "Paris is the capital"), but a fake citation appeared saying "London is the capital," the AI often changed its answer to London.
  • The Numbers: In the "General Knowledge" category, this happened 35% to 77% of the time. That means if you gave an AI a true fact and a fake reference, it was more likely to deny the truth than to stick with it.
  • The "Real Source" Effect: Even if the citation was real but the claim was false, the AI was still more likely to hallucinate than if there were no citation at all.

The Analogy: Imagine a student who knows the answer is "2+2=4." If you hand them a piece of paper that says "According to the famous Professor Smith, 2+2=5," the student stops trusting their own brain and writes down "5." The paper (the citation) overpowered the math.

What Didn't Matter?

The researchers thought certain things might make the AI trust the citation more, but they were surprised to find these factors had almost no effect:

  • Prestige: It didn't matter if the fake citation looked like it came from a top-tier university or a random blog. The AI was equally confused by both.
  • Author Identity: It didn't matter if the fake author had a name associated with a specific country or demographic. The AI didn't seem to care about who wrote it, only that something was written.
  • Model Size: You might think a "smarter" or bigger AI would be less easily tricked. But the study found that the biggest, most advanced models were just as likely to fall for the trap as the smaller ones.

The One Exception: The "Lawyer" AI

There was one domain where the AI didn't get tricked as easily: Law.

  • Why? The researchers suggest that legal text has a very specific, formal "language" (like a secret code). When the AI sees a legal citation, it seems to switch into a "scrutiny mode" and checks the facts more carefully, rather than blindly accepting the authority.
  • General Knowledge was the weakest link, where the AI was most easily fooled.

The "False Claim" Twist

The study also looked at what happens when the AI is already wrong (a false claim).

  • Some AIs got smarter: For models like Llama and Claude, seeing a citation (even a fake one) made them less likely to hallucinate on false claims. It was as if the citation made them say, "Hmm, I need to double-check this," and they corrected themselves.
  • Some AIs got dumber: For models like Gemma and Phi-4, the citation made them more likely to hallucinate on false claims. They blindly accepted the wrong info.

The Bottom Line

The paper concludes that adding a citation does not automatically make an AI more reliable. In fact, it often makes it less reliable.

If you are using an AI to help with facts, and you see it citing a source, do not assume it is telling the truth. The AI might be so impressed by the "authority" of the citation that it will happily throw away the actual truth. The study warns that for high-stakes fields (like medicine or law), we cannot assume that grounding an AI's answers in citations will fix its errors; sometimes, the citations themselves become the source of the error.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →