Can Deep Research Agents Retrieve and Organize? Evaluating the Synthesis Gap with Expert Taxonomies
This paper introduces TaxoBench, a benchmark demonstrating that current Deep Research Agents significantly lag behind human experts in both retrieving essential papers and organizing them into coherent taxonomies, revealing distinct bottlenecks in retrieval accuracy and hierarchical structure synthesis.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to write the ultimate guidebook for a massive, chaotic library. You don't just need to find the books; you need to figure out how to organize them on the shelves so that a reader can actually understand the story of the field. This is the job of a "survey," a type of academic paper that summarizes everything we know about a specific topic. For years, scientists have hoped that Artificial Intelligence (AI) could do this heavy lifting automatically. They imagined "Deep Research Agents"—super-smart AI bots that could scour the internet, find every important paper, and then arrange them into a perfect, logical map of knowledge.
But here's the catch: finding a book is easy; organizing it is hard. To understand if these AI bots are truly ready for the job, we need to test two specific skills. First, Retrieval: Can the AI find the exact same books that a human expert would choose? If the AI misses the most important papers, the guidebook is useless. Second, Organization: Once the AI has the books, can it build a "taxonomy"? Think of a taxonomy as a family tree for ideas. It's not just a list; it's a hierarchy with a main topic at the top, branching out into big categories, then smaller sub-categories, all the way down to individual papers. A good taxonomy helps you see how ideas connect. If the AI just dumps papers into a messy pile or creates a confusing, flat list, it hasn't really done the job.
This paper, titled "Can Deep Research Agents Retrieve and Organize?", puts seven of the world's most advanced AI research agents to the test. The researchers built a special challenge called TaxoBench, which uses 72 real, high-quality surveys written by human experts as the "gold standard." They asked the AI agents to either find the papers from scratch or, in a second test, organize a pre-selected list of papers into a tree structure. The results were a bit of a reality check for the AI world.
The study found that while these AI agents are getting better at talking and writing, they are still struggling with the basics of research. When asked to find the essential papers that experts use, the very best AI agent only managed to find about 20.92% of them. That means it missed nearly 80% of the most important books in the library. The other agents did even worse, with some finding fewer than 5% of the key papers. It's like asking a tour guide to show you the top 10 landmarks in a city, and they only manage to point out two or three, while missing the rest.
The second part of the test was even more revealing. Even when the researchers gave the AI agents the perfect list of papers (so they didn't have to worry about finding them), the agents still failed to organize them like a human expert. The human experts built deep, multi-layered trees with an average depth of 4.86 levels (like a tree with a trunk, big branches, smaller branches, and twigs). The AI agents, however, kept building shallow, flat structures with an average depth of only about 3 levels. They were missing the middle layers of organization.
When the researchers tried to force the AI to build deeper trees by giving it specific instructions, the AI didn't get smarter; it just got messy. It started breaking the tree apart into tiny, single-paper branches, creating a "forest" of isolated twigs instead of a structured tree. This is called "over-segmentation." The AI was so afraid of grouping things together that it ended up putting almost every paper in its own tiny, lonely category.
The paper concludes that there is a massive "synthesis gap" between what AI can do today and what human experts can do. The AI is currently stuck at a "floor" where its organizational skills are barely better than random chance. Even though the AI models are getting more powerful, they haven't closed this gap yet. The researchers suggest that before we can trust AI to write our scientific guidebooks, we need to solve two separate problems: teaching them how to find the right evidence, and teaching them how to build a deep, logical structure to hold that evidence together. Until then, human experts are still the only ones who can truly map the territory of knowledge.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.