← Latest papers
🤖 AI

Multi-Turn Evaluation of Deep Research Agents Under Process-Level Feedback

This paper introduces a multi-turn evaluation framework for deep research agents using a novel process-level feedback method called Research Gap Inference, revealing that while a single round of targeted feedback significantly improves report quality, subsequent turns fail to compound gains due to agents regressing on previously satisfied criteria, indicating that reliable multi-turn improvement remains out of reach for current architectures.

Original authors: Rishabh Sabharwal, Hongru Wang, Amos Storkey, Jeff Z. Pan

Published 2026-06-09
📖 5 min read🧠 Deep dive

Original authors: Rishabh Sabharwal, Hongru Wang, Amos Storkey, Jeff Z. Pan

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have hired a team of super-smart research assistants (called Deep Research Agents) to write a detailed report on a complex topic, like "The history of deepfake technology" or "The financial health of a specific company."

Usually, we test these assistants by giving them a question, letting them write one draft, and then grading that single draft. But in real life, if you get a draft from a human researcher, you rarely just accept it. You give them feedback like, "You missed some key laws," or "Your sources are too vague," and ask them to rewrite it.

This paper asks: Can these AI assistants actually get better when you give them feedback, or do they just make a mess of things when they try to fix their own work?

The researchers tested this using two different ways of giving feedback:

1. The "Self-Reflection" Test (The Blindfolded Editor)

In this scenario, the AI is told: "Look at your report, find your own mistakes, and fix them." No outside help is given.

  • The Result: It was like asking a student to grade their own essay without a teacher's answer key. The AI tried hard, searching the web more and reading more sources, but it couldn't figure out what was actually wrong.
  • The Analogy: Imagine you are trying to fix a leaky boat while wearing a blindfold. You might hammer the wood and add more planks (doing more work), but because you can't see the hole, you might accidentally hammer a hole in a different spot. The result? The boat is still leaking, or maybe it's leaking even more. The AI's score barely changed, or sometimes got worse.

2. The "Process-Level Feedback" Test (The Strategic Coach)

Here, the researchers didn't just say "fix this." They used a special tool called RGI (Research Gap Inference). This tool looked at the AI's report and the grading rubric to figure out how the AI failed.

Instead of saying "Add a fact about X," the feedback was more like a coach saying: "You are looking at the wrong type of sources," or "You are treating this topic too broadly and need to dive deeper into specific details," or "You missed the most important legal documents."

  • The Result: This worked wonders for the first round of fixes. The AI understood the strategy, went back to the library, found the right books, and wrote a much better report. Their scores jumped up significantly (by about 8 to 15 points).
  • The Analogy: This is like a coach telling a runner, "You aren't running out of energy; your shoes are the wrong size." Once the AI got the right shoes (the right research strategy), it ran much faster and better.

The Catch: The "Rewrite" Problem

The researchers then asked the AI to do it again (a third turn). They expected the scores to keep getting better.

  • The Result: For most of the AI models, the improvement stopped. In fact, when they tried to fix the remaining small problems, they accidentally broke the parts that were already working.
  • The Analogy: Imagine a chef who makes a great soup. The coach says, "Add a pinch of salt." The chef adds salt, and it's perfect. But then the coach says, "Now, fix the garnish." The chef, trying to be perfect, decides to throw out the whole pot and start from scratch. In the process of making a new pot, they forget to add the salt they just added, or they mess up the vegetables they got right the first time.
  • Why? The AI models used in this study work by "rewriting the whole report from scratch" every time. They don't have a memory of "what was good last time." So, when they try to fix one thing, they often lose the things they had already fixed.

The One Exception: The "Careful" Model

One specific AI model (DeepSeek-V4-Flash) behaved differently. It didn't throw out the whole pot. It kept most of the good soup and just added the new ingredients. It improved its score again in the third round.

  • The Cost: However, this "careful" model was incredibly expensive and slow. It used about 4 times more computer power and took much longer to think than the others. It was like a chef who refuses to throw anything away and meticulously checks every single ingredient, which takes forever and costs a fortune.

The Bottom Line

  1. AI can't fix itself: If you tell an AI to "fix your own mistakes," it usually just makes things worse or stays the same because it can't see the big picture.
  2. Strategic feedback works: If you tell the AI how to research (e.g., "look for official documents, not blog posts"), it can write a much better report.
  3. The "Full Rewrite" flaw: Current AI research tools are like artists who erase their entire canvas every time they want to add a new brushstroke. This makes it very hard to improve a report step-by-step without accidentally deleting the good parts. To get truly reliable, multi-step improvements, we need AI that can remember what it got right and only change what needs fixing, rather than starting over every time.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →