← Latest papers
🤖 AI

Benchmarking Agentic Review Systems

This paper benchmarks several agentic review systems against human quality judgments and injected errors, demonstrating that while they still require improvement, top configurations like OpenAIReview with GPT-5.5 effectively track paper quality, detect a majority of errors, and receive positive feedback from real users.

Original authors: Dang Nguyen, Wanqing Hao, Yanai Elazar, Chenhao Tan

Published 2026-06-19
📖 5 min read🧠 Deep dive

Original authors: Dang Nguyen, Wanqing Hao, Yanai Elazar, Chenhao Tan

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the world of academic research as a massive, chaotic library where thousands of new books (research papers) are added every day. Traditionally, a small team of human librarians (peer reviewers) tries to read every single book to decide which ones are masterpieces and which are full of holes. But now, AI is helping people write books faster than ever, and the librarians are drowning. They are so overwhelmed that they might start missing big mistakes or accepting bad books just to clear the queue.

This paper is like a report card for a new class of "AI Librarians" designed to help the humans. The researchers built a testing ground to see if these AI systems can actually do a good job, or if they are just making noise.

Here is what they found, explained through simple analogies:

1. The "Volume Test": Do they know a bad book when they see it?

The researchers asked a simple question: If an AI reads a terrible paper, will it complain more than if it reads a great paper?

They tested this on real papers from top conferences. They found that the AI systems do seem to sense quality.

  • The Analogy: Imagine a food critic. If you give them a burnt steak, they write a long, angry review. If you give them a perfect steak, they write a short, happy note. The AI critics in this study acted the same way: they wrote significantly more comments on the "bad" papers than the "good" ones.
  • The Winner: The best combination was a system called OpenAIReview powered by a very smart model called GPT-5.5. It was about 83% accurate at spotting the difference between a high-quality paper and a low-quality one just by how much it complained.

2. The "Spot the Difference" Game: Can they catch specific errors?

Knowing a paper is "bad" isn't enough; the AI needs to find the actual mistakes. To test this, the researchers played a trick on the AI. They took clean, perfect papers and secretly injected four types of errors:

  • Math mistakes: Changing a number or a sign in an equation.
  • False claims: Lying about what a study proved.
  • Bad logic: Making a silly argument that doesn't follow.
  • Flawed experiments: Designing a test that can't work.

They then asked the AI to find these "hidden traps."

  • The Result: The best AI system caught about 71.6% of the traps. That's pretty good, but it means it still missed nearly 3 out of 10 errors.
  • The "Swarm" Effect: Here is the cool part. Different AI models found different errors. One model might catch a math error, while another catches a logic error. When the researchers combined the findings of all six different AI models they tested, they caught 83.3% of the errors.
  • The Metaphor: It's like having a team of detectives. Detective A is great at finding fingerprints, but Detective B is better at finding footprints. If you use just one detective, you miss clues. If you use the whole team, you catch almost everything.

3. The "Real World" Test: Do actual authors like them?

The researchers didn't just test the AI in a lab; they put OpenAIReview on a public website and let real researchers use it on their own papers. They collected feedback from over 1,000 papers.

  • The Verdict: The users liked the AI! For every 1 person who disliked a comment, 1.44 people liked it. Many users even marked the comments as "fixed," meaning they actually went back and changed their papers based on the AI's advice.
  • The Complaint: The main reason people disliked the AI wasn't that it missed big errors; it was that the AI was too picky. It would flag tiny, unimportant things (like a minor formatting quirk) or point out things that weren't actually mistakes. It was like a strict editor who corrects your spelling but also complains that you used the word "very" too many times.

The Big Takeaway

The paper concludes that AI reviewers are ready to be helpful partners, but they aren't perfect replacements yet.

  • They are good at: Sensing the overall quality of a paper and finding specific technical errors (especially when you use a "team" of different AI models).
  • They are bad at: Being precise enough to avoid nagging about tiny, unimportant details.

The authors suggest that the future isn't about replacing human reviewers, but about using AI as a safety net. The AI can do the boring, exhaustive work of checking every equation and claim for errors, while the human reviewers can focus on the big picture: "Is this idea actually new and important?"

In short: The AI librarians are smart enough to spot the bad books and find the typos, but they still need a human to tell them when they are being too annoying about the font size.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →