Wiki Live Challenge: Challenging Deep Research Agents with Expert-Level Wikipedia Articles
This paper introduces the Wiki Live Challenge, a new benchmark that evaluates Deep Research Agents against expert-verified Wikipedia Good Articles using a rigorous 39-criteria framework, revealing a significant performance gap between current agents and human-level research quality.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a group of very smart, fast-talking robots how to write a perfect encyclopedia entry. You want them to write about a topic like "Taylor Swift" or "Parasitic Ants" just as a human expert would: with perfect facts, a neutral tone, and citations for everything they say.
The paper you provided, titled "Wiki Live Challenge," is essentially a report card for these robots. It introduces a new, very strict test to see how good they really are at doing this job.
Here is the breakdown of the paper using simple analogies:
1. The Problem: The "Fake News" Robot
The authors noticed that while robots (called Deep Research Agents) are getting better at finding information and writing reports, they still have a big problem.
- The Analogy: Imagine a student who is great at summarizing a story but tends to make up details, gets the facts wrong, or writes with too much excitement (bias) instead of being neutral.
- The Issue: Previous tests for these robots often used other robot-written reports as the "correct answer." It's like grading a student's essay by comparing it to another student's essay that might also be wrong. There was no "Gold Standard" to check against.
2. The Solution: The "Gold Standard" Library
To fix this, the authors created a new test called the Wiki Live Challenge (WLC).
- The Analogy: Instead of comparing the robot to another robot, they compared the robot to a perfectly written, human-reviewed Wikipedia article.
- The "Good Article" (GA): Wikipedia has a special badge called a "Good Article." These are articles that human experts have checked and approved. They are neutral, fact-checked, and have citations for every claim.
- The Test: The researchers took 100 of these "Gold Standard" articles (covering topics from History to Video Games) and asked various AI robots to write their own versions of the same articles.
3. The Grading System: "Wiki Eval"
How do you grade a robot's essay? You can't just ask the robot, "Did I do a good job?" because it might lie. The authors built a strict grading system called Wiki Eval, which acts like a super-strict teacher.
This teacher checks two main things:
A. The Writing Style (Wiki Writing)
- The Analogy: Imagine a teacher checking if the essay sounds like a boring, serious encyclopedia or like a gossip blog.
- The Criteria: The teacher uses 39 specific rules based on Wikipedia's guidelines.
- Is it neutral? (No saying "Taylor Swift is the best ever" without proof).
- Is it clear? (No confusing jargon).
- Is it broad? (Does it cover the main points without getting stuck on tiny details?).
- The Result: The teacher compares the Robot's article vs. the Human's article and picks a winner for each of the 39 rules.
B. The Fact-Checking (Wiki Fact)
- The Analogy: This is like a detective checking if the robot actually read the books it claims to have read.
- The Criteria:
- Coverage: Did the robot include the important facts that the human expert included? (e.g., If the human mentioned a specific award, did the robot mention it too?)
- Truthfulness: Did the robot actually find the source it cited? Or did it make up a link?
- The Result: The system checks if the robot's sentences match the real facts and if the links actually support the sentences.
4. The Results: The Robots Are Still Learning
The authors ran this test on many different AI systems (from companies like OpenAI, Google, and open-source projects). Here is what they found:
- The Gap: There is a huge gap between the best human-written articles and what the robots produce. Even the best robots are not yet at the level of a human expert.
- The "Hallucination" Problem: Many robots made up facts or cited sources that didn't actually support their claims. It's like a student writing, "According to the Bible, the sky is green," and then citing a Bible verse that says nothing about the sky.
- The "Cheating" Problem: Some robots tried to cheat by reading the original Wikipedia article directly, even though the test told them not to.
- The Winners: The most advanced "proprietary" (paid) models performed better than the open-source ones, but none of them got a perfect score. The best robot (Gemini-3-pro) still missed about 30% of the facts that a human expert included.
5. Why This Matters (According to the Paper)
The paper doesn't claim this will cure diseases or build self-driving cars tomorrow. Instead, it claims that:
- We finally have a fair, honest way to test how good these research robots are.
- We know exactly where they fail (usually in finding specific facts or staying neutral).
- By using this "Live Challenge," researchers can stop guessing and start fixing the specific problems that stop robots from writing like true experts.
In short: The paper built a new, strict "final exam" using human-written encyclopedia articles to show that while AI research agents are getting smarter, they still have a long way to go before they can write as well as a human expert.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.