← Latest papers
🤖 AI

Why Your Deep Research Agent Fails? On Hallucination Evaluation in Full Research Trajectory

This paper addresses the limitations of outcome-based evaluation in Deep Research Agents by introducing the PING Taxonomy and the DeepHalluBench benchmark to enable process-aware, fine-grained auditing of hallucinations across the full research trajectory, revealing systemic reliability gaps and offering actionable insights for architectural improvements.

Original authors: Yuhao Zhan, Tianyu Fan, Linxuan Huang, Zirui Guo, Chao Huang

Published 2026-05-26
📖 5 min read🧠 Deep dive

Original authors: Yuhao Zhan, Tianyu Fan, Linxuan Huang, Zirui Guo, Chao Huang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you hire a super-smart research assistant to find the answer to a very tricky question. You ask them, "Who is the only person to win an Oscar for acting in a movie they also directed?" They go off, search the internet, read articles, and come back with a final report.

The Old Way of Checking Their Work:
Previously, if you wanted to know if your assistant was good, you just looked at the final report. If the name was right, they got a gold star. If the name was wrong, they got a red X.

The Problem:
This paper argues that looking only at the final answer is like judging a chef only by the taste of the soup, without watching how they cooked it. Maybe the soup tasted bad because they forgot the salt (a small mistake), or maybe they used a rotten tomato they found in the back of the fridge (a big lie), or maybe they started cooking the wrong dish entirely because they misunderstood your order.

The authors say that current "Deep Research Agents" (AI researchers) are making mistakes during the process, but because we only check the final answer, we don't see where or why they failed. The mistakes hide in the middle.

The New Approach: The "Process Detective"

The authors built a new system called DeepHalluBench. Think of this as a transparent kitchen where we can watch the assistant cook step-by-step. They don't just check the soup; they check every ingredient, every chop, and every stir.

They created a new rulebook called the PING Taxonomy to categorize the different ways the AI gets confused. Here is what PING stands for, using simple analogies:

  1. P - Propagation (The "Domino Effect"):
    Imagine the assistant makes a small mistake in step one, like thinking "Apples are blue." In step two, they use that wrong idea to search for "Blue Apples." In step three, they write a report about "Blue Apples." The whole chain of events is built on that first lie. This is Propagation. The paper found that once an AI starts lying, it often keeps building on that lie, making the final report completely wrong even if the later steps were done "correctly" based on the wrong info.

  2. I - Intent (The "Misunderstood Order"):
    You asked for a "vegan" recipe, but the assistant ignored that word and gave you a steak recipe. Or you asked for a "2023" report, and they gave you a "2010" one. The assistant did the research, but they didn't listen to your specific rules. This is Intent hallucination.

  3. N - Noise-induced (The "Distracted Librarian"):
    Imagine the assistant goes to the library and finds 100 books. 99 of them are about cats, but 1 is about the specific dog you asked about. The assistant reads all 99 cat books and ignores the one dog book because it's buried in the pile. They didn't lie, but they missed the most important clue. This is Noise-induced hallucination.

  4. G - Grounding (The "Fake Fact"):
    The assistant looks at a real article but then says, "This article says the moon is made of cheese." The article actually says nothing about cheese. The assistant is making things up or blaming the wrong source. This is Grounding hallucination.

What They Found (The Results)

The authors tested six famous AI research assistants (like those from OpenAI, Google, and others) using their new "Process Detective" system on 100 very hard questions.

  • Everyone is struggling: Even the best AI assistants made mistakes in the middle of their research. None of them were perfect.
  • The "Domino" is the biggest problem: The most dangerous issue wasn't just making one lie; it was that the AI would make a small mistake early on, and then the rest of the research would be built on that mistake, leading to a total failure.
  • They get stuck on the first thing they see: The AI tends to focus heavily on the first few search results it finds (like a person who only reads the first page of a book and assumes they know the whole story). If the answer is on page 50, the AI often misses it.
  • They prefer "boring" answers: If the AI finds five articles that all say the same thing, it loves them. If it finds one unique article that says something different, it often ignores it.

The Conclusion

The paper concludes that simply making the AI "smarter" or giving it more internet access isn't enough. The problem is how the AI thinks while it searches.

To fix this, we need to build AI that can:

  1. Catch its own mistakes early (before they turn into a domino effect).
  2. Listen better to the specific rules you give it.
  3. Not get distracted by the first thing it sees, but keep looking for the best answer.

In short: We can't just grade the final report anymore. We have to watch the whole movie to understand why the AI failed.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →