Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents
This paper introduces a novel evaluation framework that uses an AST parser to assess LLM-generated research reports across link validity, relevance, and factual accuracy, revealing a critical disconnect where even frontier models maintain high citation accessibility and relevance but struggle with factual consistency, a problem that worsens as retrieval depth increases.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are hiring a team of super-smart research assistants (AI models) to write a report on a complex topic. You tell them, "Go find the truth on the internet, write a report, and make sure to list exactly where you found every single fact."
The paper you shared is like a quality control inspector who decided to check the homework these AI assistants turned in. Instead of just reading the report and nodding along, the inspector built a special robot to check every single footnote the AI wrote.
Here is what they found, explained simply:
The Three-Step Inspection
The researchers built a system to check the AI's citations in three specific ways, like a three-part security check:
- The "Is the Door Open?" Check (Link Works): Does the link actually work? If you click it, does it take you to a webpage, or is it a broken 404 error?
- The "Is This About the Same Thing?" Check (Relevant Content): If the link works, does the webpage actually talk about the topic the AI claimed it did? (e.g., The AI says "This link proves cats are dogs," but the link is actually a recipe for lasagna).
- The "Is It True?" Check (Fact Check): This is the hardest part. Does the webpage actually contain the specific facts, numbers, or dates the AI said it did?
The Big Surprise: The "Shiny Wrapper" Problem
The most shocking discovery is that the AI assistants are great at the first two checks but terrible at the third.
- The Analogy: Imagine a student who writes a paper and includes a bibliography.
- 94% of the time, the books in the bibliography actually exist on the library shelves (Link Works).
- 80% of the time, those books are actually about the subject the student is writing about (Relevant Content).
- But, when you actually open the book and read the specific page the student cited, only about 40% to 77% of the time does the book actually say what the student claimed it said.
The Takeaway: The AI is very good at finding a working link to a relevant website, but it is often lying about what that website says. It's like a tour guide who always takes you to the right museum, but then tells you the wrong story about the exhibits inside.
The "More is Less" Paradox
The researchers also tested what happens if they tell the AI to search deeper and look at more websites. You might think, "If I ask the AI to read 150 websites instead of just 2, it will be more accurate, right?"
Wrong.
- The Analogy: Imagine trying to solve a puzzle. If you have 2 puzzle pieces, you can easily see the picture. If someone dumps 150 puzzle pieces in front of you, you get overwhelmed. You might start forcing pieces together that don't fit just to finish the picture.
- The Result: As the AI was forced to look at more and more sources (from 2 up to 150), its ability to tell the truth actually dropped by about 42%. The links still worked, and the topics were still relevant, but the specific facts became a mess. The AI got "information overload" and started hallucinating details.
Who Did Well? Who Did Poorly?
- The "Frontier" Models (The Big Tech AIs): These models (like the ones from OpenAI, Anthropic, and Google) were very good at generating reports that looked perfect. They almost always found working links and relevant topics. However, even the smartest ones only got the facts right about 39% to 77% of the time.
- The Open-Source Models: These models struggled even more. Many of them couldn't even finish the task of writing a report with citations. When they did try, they were much worse at finding working links and getting facts right.
The Bottom Line
The paper concludes that we cannot trust AI research just because it has a list of links at the bottom. A long list of working links creates a "false sense of trust." Just because the door is open and the room is the right one, doesn't mean the story the AI tells you about what's inside the room is true.
The researchers built this "inspection robot" to help us see that gap between a citation that looks good and a citation that is actually true.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.