Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents
This paper systematically evaluates the validity of citation URLs in large language models and deep research agents, revealing significant rates of hallucination and non-resolving links across domains, and demonstrates that an open-source tool called `urlhealth` can effectively reduce these errors by up to 79 times through agentic self-correction.
Original paper dedicated to the public domain under CC0 1.0 (http://creativecommons.org/publicdomain/zero/1.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are asking a very smart, well-read librarian (an AI) to write a research paper for you. You ask, "What are the latest breakthroughs in cancer research?" The librarian hands you a beautiful, detailed report. But here's the catch: at the bottom of every paragraph, there are little footnotes pointing to specific books or articles on the shelves that supposedly prove the librarian's claims.
This paper is like a team of detectives who went to the library to check if those footnotes actually lead to real books, or if the librarian just made them up.
Here is the story of their investigation, broken down simply:
1. The Problem: "Ghost Books"
The researchers found that while these AI librarians are amazing at writing, they sometimes suffer from "citation hallucinations." This means they invent URLs (web addresses) that look real but lead nowhere.
- The Analogy: It's like the librarian pointing to a book on the shelf and saying, "Read page 42 of The History of Time Travel," but when you go to pull that book off the shelf, it's not there. The book never existed.
- The Reality: They found that between 3% and 13% of the links these AIs provide are "ghosts"—they never existed in the first place. Another 5% to 18% are "dead links"—the book used to be there, but someone took it off the shelf (link rot).
2. The Suspects: "Deep Researchers" vs. "Quick Searchers"
The study looked at two types of AI:
- The Quick Searchers: These are AIs that do a standard Google search and give you an answer.
- The Deep Researchers: These are the new, fancy AIs that spend a lot of time "thinking," reading multiple pages, and writing long, detailed reports with hundreds of footnotes.
The Surprise: You might think the "Deep Researchers" are more careful because they work harder. But the study found the opposite! Because they try to write so much and cite so many sources, they actually make up more fake links than the quick searchers. It's like a student who tries to write a 50-page essay in one night; they are more likely to invent facts just to fill the pages.
3. The "Subject" Effect: Some Topics Are Trickier
The researchers noticed that the AIs struggle more with certain topics.
- The Analogy: Imagine asking the librarian about "Business" (stable, predictable) vs. "Theology" or "Medicine" (fast-changing, complex).
- The Result: The AIs were much more likely to invent fake links when talking about Theology or Medicine. Why? Because these fields are full of specialized, hard-to-find, or rapidly changing information. The AI gets confused and starts guessing. In contrast, for "Business" or "Architecture," the links were much more reliable.
4. The "More is Not Better" Rule
There was a myth that if an AI gives you more citations, the quality must be better. The study proved this wrong.
- The Analogy: If a chef gives you a plate with 100 ingredients, it doesn't mean the meal is better. It just means there's a higher chance of finding a rotten tomato in the pile.
- The Result: AIs that generated the most citations actually had the highest error rates. Quantity did not equal quality.
5. The Solution: The "URL Health Check" Tool
The researchers didn't just point out the problem; they built a tool to fix it. They created an open-source tool called urlhealth.
- How it works: Imagine a "spell-checker" for links. Before you trust the AI's report, this tool runs a quick check on every single link.
- Is the link alive? (Green light)
- Is it a ghost? (Red light - never existed)
- Is it dead? (Yellow light - used to exist, now gone)
- The Magic: When they let the AI use this tool to "self-correct" (check its own work), the number of fake links dropped dramatically—by 6 to 79 times! The AI went from having a messy, unreliable report to a clean, trustworthy one.
The Big Takeaway
This paper teaches us two main things:
- Don't trust everything you read: Even the smartest AI can make up fake sources, especially when it's trying to be too helpful or too detailed.
- Verification is key: We need tools that automatically check if a link is real before we trust the information. Just like you wouldn't buy a house without checking the deed, you shouldn't trust an AI report without checking the links.
The researchers have made their "URL Health Check" tool available for everyone to use, hoping to stop the spread of fake citations in the future.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.