A Benchmark for Deep Information Synthesis
The paper introduces DEEPSYNTH, a challenging new benchmark comprising 120 real-world tasks across 7 domains that evaluates the ability of LLM-based agents to synthesize information and infer insights, revealing significant current limitations in reasoning and hallucination control.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are hiring a new research assistant to help you plan a massive, complex international trip. You don't just want them to Google "best hotels in Paris." You want them to:
- Check flight prices in Tokyo, currency exchange rates in Berlin, and visa laws in Brazil.
- Read three different government reports on tourism trends.
- Cross-reference a spreadsheet of airline delays with a news article about a strike.
- Finally, write a single, perfect itinerary that explains why you should go to Country X instead of Country Y, backed by hard data.
This is exactly what the paper DEEPSYNTH is about. It's a new "exam" designed to test how good AI agents (smart computer programs) are at doing this kind of heavy lifting.
Here is the breakdown of the paper using simple analogies:
1. The Problem: The "Google Search" Trap
Right now, most AI tests are like asking a student, "Who was the first president of the US?" The AI just remembers the answer from its training data. It's a simple fact retrieval.
But real life isn't like that. Real life is messy. It's like asking, "Which non-ASEAN countries recovered their tourism to Singapore by 2023, and was it mostly business or vacation travelers?"
- The Challenge: The AI can't just "remember" this. It has to go out, find data from 67 different countries, read tables, compare numbers, and do math.
- The Current State: The paper says current AIs are like students who are great at reciting the dictionary but terrible at writing a thesis. They get lost when they have to connect dots between different sources.
2. The Solution: The "DEEPSYNTH" Exam
The authors created a new benchmark (a test) called DEEPSYNTH.
- The Content: It has 120 very hard questions.
- The Difficulty: To answer just one question, the AI has to visit an average of 4.2 different websites, read up to 15 documents, and perform complex math or logic.
- The "Gold Standard": Humans (experts) created these questions. They didn't just ask a question; they first found the data, did the analysis, figured out the answer, and then wrote the question. This ensures the answer isn't something the AI could have just memorized from the internet before.
Analogy: Imagine a cooking contest.
- Old Tests: "What is the boiling point of water?" (Easy, memorized).
- DEEPSYNTH: "Here are 10 different recipes from 10 different countries, a list of ingredients available in your local market, and a budget. Create a menu that balances nutrition and cost, and explain your choices."
3. The Results: The AI Struggles
The authors tested the smartest AI models available (like GPT-5, DeepSeek, and specialized "research agents") on this exam.
- The Score: The results were shocking. The best AI only got an F1 score of 8.97 (out of 100).
- The Reality Check: Under the strictest rules (Exact Match), almost every AI got a 0. They couldn't get a single answer perfectly right.
- Why?
- Hallucinations: The AI makes things up. It might invent a statistic because it "thinks" it sounds right.
- Navigation Errors: It clicks the wrong link or gets stuck on a website.
- The "Africa" Gap: The AI was terrible at answering questions about African countries. This suggests the AI is biased because it was trained mostly on data from the US, Europe, and Asia.
Analogy: It's like giving a brilliant librarian a map to a library with 10,000 books. The librarian knows how to read, but when asked to find a specific fact hidden in a book in the "Africa" section, they either can't find the section, or they guess the answer because they've never seen that section before.
4. The "Deep Research" Agents
The paper also tested "Agents"—AIs that are programmed to use tools like web browsers and code calculators.
- The Result: Even these "super-agents" only solved 3 out of 120 tasks perfectly.
- The Lesson: Just giving an AI a web browser doesn't make it smart enough to do deep research. It still gets confused when it has to plan a long chain of steps (e.g., "Go to site A, download the file, open it, find the number, go to site B, compare it...").
5. Why This Matters
The paper concludes that we are hitting a wall. We can't just make AI "smarter" by making it bigger; we need to teach it how to synthesize information.
- Synthesis is the ability to take a pile of raw, messy data and turn it into a clear, useful insight.
- Currently, AI is like a very fast photocopier. It can copy facts, but it can't yet write the story that connects them.
Summary
DEEPSYNTH is a reality check for the AI industry. It shows that while AI is getting better at chatting and writing, it is still very bad at doing the kind of deep, multi-step research that humans do every day. The AI is currently a "super-fast fact-checker" that needs to become a "true researcher" before it can be trusted with complex real-world problems.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.