← Latest papers
🔬 physics

LLM hallucinations in the wild: Large-scale evidence from non-existent citations

This study provides large-scale evidence that the widespread adoption of large language models has led to a sharp rise in non-existent citations across millions of scientific papers, threatening the reliability and equity of knowledge production while outpacing current editorial safeguards.

Original authors: Zhenyue Zhao, Yihe Wang, Toby Stuart, Mathijs De Vaan, Paul Ginsparg, Yian Yin

Published 2026-05-11
📖 5 min read🧠 Deep dive

Original authors: Zhenyue Zhao, Yihe Wang, Toby Stuart, Mathijs De Vaan, Paul Ginsparg, Yian Yin

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: The "Fake Book" Problem

Imagine you are writing a school essay. You need to list the books you read in your bibliography. If you make up a book title that doesn't exist, your teacher can easily check and say, "This book isn't real."

This paper is a massive investigation into what happens when AI (Large Language Models) starts writing those bibliographies for scientists. The researchers found that AI is increasingly inventing fake book titles (citations) and slipping them into real scientific papers. While scientists have always made small mistakes, the study shows that since AI became popular, the number of these "fake books" has exploded.

How They Caught the AI

The researchers treated the scientific world like a giant library. They looked at 111 million references (citations) across four major digital libraries (arXiv, bioRxiv, SSRN, and PubMed Central).

  • The Method: They built a digital "search engine" to check every single citation. If the search engine couldn't find the book or paper in the real world, they flagged it as "unmatched."
  • The Filter: They knew some citations were just messy (bad formatting or obscure books). So, they used a smart AI to clean up the messy ones. If a citation still couldn't be found after all the cleaning, they counted it as a hallucination (a fake reference).
  • The Baseline: They compared the number of fake citations before AI was popular (2022) to after. They found a sharp, scary spike starting in 2024.

The Result: In 2025 alone, they estimate there are 146,932 fake citations just in these four libraries. And that's likely just the tip of the iceberg.

Who Is Doing This?

The paper found that the "fake book" problem isn't coming from a few bad apples; it's spread out everywhere, but it's concentrated in specific groups:

  1. The "New Kids on the Block": The scientists most likely to include fake citations are early-career researchers (students and new professors) who haven't published many papers yet.
    • Analogy: Think of it like a new driver who hasn't learned the rules of the road yet. They are using AI to help them drive (write papers), but they are accidentally driving into fake neighborhoods because they don't know the map well enough to spot the errors.
  2. The "Solo Drivers": Papers written by single authors or small teams have much higher rates of fake citations than big teams.
    • Analogy: If you write a story alone, you might miss a typo. If you write it with a team of editors, someone else will catch it. Big teams act as a safety net; small teams often don't have that net.
  3. The "AI-Heavy" Fields: Fields where people use AI the most (like Computer Science and Social Sciences) have the highest rates of fake citations.

Who Gets the Credit?

Here is the unfair part. When AI invents a fake paper, it doesn't just make up a random name. It tends to invent papers written by famous, successful, male scientists.

  • Analogy: Imagine a student making up a fake quote for a history essay. Instead of making up a random name, the student accidentally (or subconsciously) makes up a quote from the most famous president in history.
  • The Consequence: This means AI is accidentally reinforcing inequality. It gives extra "credit" to people who are already famous and male, while ignoring less famous or female scholars. Even if the paper is fake, the famous person's name gets added to the list of "people who wrote about this."

Why Don't Editors Catch It?

You might think, "Don't journals check these papers before publishing them?" The study says no, not really.

  • The Preprint Problem: Many scientists post their work online before it's officially published (like a draft on a blog). The study found that 78.8% of the fake citations slip past the initial online moderators.
  • The Peer Review Gap: Even when these papers go to formal journals for review, 85% of the fake citations survive the process and get published.
  • Analogy: It's like a security guard at a concert who checks for tickets. The guard catches the people with really obvious fake tickets, but the people with "good-looking" fake tickets (which AI makes very convincing) walk right through.

The "Snowball" Effect

The most dangerous part of this is how the fake citations spread.

  1. The Library Gets Contaminated: Once a fake paper is published, it gets added to Google Scholar and other databases.
  2. The Loop: Future researchers (and future AIs) will search these databases, see the fake paper, and think, "Oh, this is a real source!" They will then cite it in their own work.
  3. The Result: The fake paper becomes "real" because everyone treats it as real. The AI is learning from its own mistakes, creating a cycle where the library fills up with more and more fiction.

The Bottom Line

This paper argues that we are in a crisis of trust. Science relies on the idea that if you cite a paper, that paper actually exists and says what you claim it says.

The study shows that AI is breaking this trust. Even in science—the field with the strictest rules and the best tools for checking facts—AI is successfully sneaking in thousands of fake references every year. If this is happening in science (where it's easiest to catch), the paper suggests it is likely happening even worse in other fields like law, medicine, and government reports, where there are fewer tools to check the facts.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →