← Latest papers
💬 NLP

One Interaction Is Worth a Thousand Guesses: Benchmarking the Interactive Capabilities of Deep Research Agents

This paper introduces IDRBench, the first benchmark designed to systematically evaluate the interactive capabilities of deep research agents by measuring how effectively they solicit clarification and align with evolving user intent, demonstrating that such interaction significantly improves research quality and robustness across various large language models.

Original authors: Yingchaojie Feng, Qiang Huang, Xiaoya Xie, Zhaorui Yang, Jun Yu, Wei Chen, Anthony K. H. Tung

Published 2026-06-23
📖 4 min read☕ Coffee break read

Original authors: Yingchaojie Feng, Qiang Huang, Xiaoya Xie, Zhaorui Yang, Jun Yu, Wei Chen, Anthony K. H. Tung

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you hire a brilliant, super-fast research assistant to write a report for you.

The Old Way (Autonomous Agents):
You give them a vague instruction like, "Tell me about the history of coffee." They immediately disappear into a library, start reading, and come back hours later with a 50-page essay.

  • The Problem: If you actually wanted to know about coffee farming in Ethiopia but they spent all that time writing about espresso machines in Italy, you're stuck with a report you can't use. They guessed your intent, and they guessed wrong. You can't stop them mid-process to say, "Wait, I meant something else!"

The New Idea (Interactive Agents):
This paper introduces a new way of working where the assistant doesn't just guess. Instead, they pause and ask you questions.

  • The Analogy: It's like a detective working on a case. Instead of running off to solve the whole mystery alone, they call you every hour to say, "I found a clue about the butler, but I'm not sure if you want me to focus on the butler or the gardener. Which one should I investigate next?"
  • The Result: By asking these questions, the assistant stays on the right track, saving time and making sure the final report is exactly what you wanted.

What is IDRBench?

The authors built a testing ground called IDRBench (Interactive Deep Research Benchmark). Think of it as a giant, controlled "training gym" for these AI assistants.

  1. The Setup: They took 100 real-world research topics but deliberately made the instructions vague and incomplete (like telling a chef to "make a meal" without saying what ingredients you have or what you're allergic to).
  2. The Test: They watched to see which AI assistants were brave enough to stop and ask, "What kind of meal did you have in mind?" versus which ones just blindly started cooking and hoped for the best.
  3. The Simulator: Since they couldn't ask 100 real humans to play along, they built a "Robot User" (a User Simulator). This robot acts like a real person, giving feedback based on a hidden "answer key" (the reference document) without actually cheating by showing the answer.

What Did They Find?

They tested seven different powerful AI models (some from big companies, some open-source) in two modes: Solo Mode (no talking) and Chat Mode (asking questions).

  • Talking Always Helps: Every single AI model got better at its job when allowed to ask questions. The final reports were more accurate, covered the right topics, and felt more "human."
  • The "Weak" Models Benefit Most: The AI models that were a bit less smart on their own improved the most when they were allowed to chat. It's like a student who struggles with math but gets an A when they are allowed to ask the teacher for a hint.
  • The "Strong" Models Still Win: Even the smartest models (like GPT-5.1) got better with a little chat, but they didn't need as many questions to get there.
  • Cost vs. Benefit: Asking questions takes a little bit of extra time and money (computing power). However, the paper found that for most models, the extra cost was tiny compared to the huge jump in quality. One model even saved money by asking questions because it stopped wasting time researching the wrong things!

The Bottom Line

The paper argues that the future of AI research isn't about building assistants that are perfect at guessing what you want. It's about building assistants that are good at talking to you.

Just like a human expert would ask for clarification before starting a big project, the best AI agents are the ones that know when to pause, ask a question, and make sure they are on the same page before doing the heavy lifting. The authors created IDRBench to prove that this "interactive" skill is just as important as raw intelligence.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →