Evaluating Deep Research Agents on Expert Consulting Work: A Benchmark with Verifiers, Rubrics, and Cognitive Traps
This paper introduces a rigorous benchmark for evaluating deep research agents on expert consulting tasks using SME-authored prompts with cognitive traps, revealing that current frontier models (Claude, o3, and Gemini) achieve uniformly low acceptance rates due to distinct failure modes like hallucinations, arithmetic errors, and inconsistent reasoning.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are hiring a team of super-smart research assistants to do a complex job: they need to read a messy pile of documents, find specific numbers, do some math, and write a professional report for a CEO. If they get even one number wrong, the CEO might make a billion-dollar mistake.
This paper is like a rigorous job interview for three of the most advanced AI research assistants available today (Claude, o3, and Gemini). The authors, from Deccan AI Research, wanted to see if these AIs could actually do the kind of detailed, high-stakes work that human management consultants do, or if they were just good at answering simple trivia questions.
Here is the breakdown of their experiment, using simple analogies:
1. The Test: A "Trap-Filled" Obstacle Course
Most previous tests for AI were like asking, "What is the capital of France?" (Factual recall). This test was different. The authors created 42 complex scenarios based on real consulting work.
Think of these scenarios as a survival course designed to trick the AIs. The documents given to the AIs contained "cognitive traps," such as:
- The "Red Herring" Trap: A file full of 100 pages of text, but the answer is hidden in a tiny footnote on page 98.
- The "Contradiction" Trap: One document says the price is $50, but a footnote says, "Unless it's Tuesday, then it's $60."
- The "Math Trap": Asking for an average, but the correct method requires a weighted average (like calculating your grade based on how much each test counts).
If an AI just skimmed the surface or guessed, it would fail.
2. The Judges: Two Layers of Grading
The paper didn't just let an AI grade the AI. They used a two-layer scoring system:
- Layer 1: The Robot Checkers (Verifiers): These are strict, automated checks. Did the AI produce the file? Did it use the right numbers? Did it cite the right sources? If the AI got a single math step wrong, this layer marked it as a failure.
- Layer 2: The Human Expert (The SME): A real human consultant (with 15+ years of experience) read the AI's report. They graded it on five things:
- Data Integrity: Did you make up facts?
- Analytical Rigor: Is your logic sound?
- Relevance: Did you answer the question or just ramble?
- Execution: Did you do the math right?
- Format: Is the report professional and usable?
The final score combined these two layers. To "pass" the test, an AI had to clear both the robot checkers and impress the human expert.
3. The Results: The "Pass Rate" Was Shockingly Low
Out of 126 total attempts (42 tasks × 3 AIs), almost no one passed.
- Gemini passed about 21% of the time.
- Claude and o3 passed only 9.5% of the time.
In other words, if you hired one of these AIs for a million-dollar consulting project, there was a 79% to 90% chance they would fail to deliver a usable, accurate report.
4. How Each AI Failed (The "Personalities" of Failure)
The paper found that each AI failed in its own unique way, like three different students taking a hard exam:
Claude (The Confident Fabricator):
- Strength: It was the best at actually creating the final files (like Word docs or Excel sheets). It rarely forgot to turn in the assignment.
- Weakness: It was the most likely to lie. When it couldn't find an answer in the documents, it would confidently invent a number or a source to fill the gap. It was like a student who wrote a beautiful essay but made up all the facts.
o3 (The Careful but Clumsy Mathematician):
- Strength: When it did reason, its logic was often the cleanest.
- Weakness: It frequently dropped required sections of the report or made cascading math errors. If it messed up step one, the rest of the math was wrong. It also sometimes ignored instructions to create a file and just gave a text answer instead.
Gemini (The Rollercoaster):
- Strength: When it worked, it was often the best. It had the highest number of "perfect" scores.
- Weakness: It was unreliable. It had the highest number of "zero" scores. Sometimes it would crash, refuse to answer, or produce a file with broken code. It was like a student who either aced the test or failed so badly they didn't even show up.
5. The Big Takeaway
The paper concludes that while these AI agents are impressive, they are not yet ready for high-stakes, decision-grade work.
If a human consultant makes a mistake, it's usually a small slip. If these AIs make a mistake, it's often a catastrophic failure: they might invent a fake source, crash the software, or miss a critical footnote that changes the entire business decision.
The authors are currently preparing a second, larger version of this test, but for now, the message is clear: Don't trust these AIs with your company's billion-dollar decisions yet. They are still learning how to be reliable researchers.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.