← Latest papers
💬 NLP

How Closely Do LLM Reviews Align with Human Peer Review?

This study evaluates three leading large language models against human peer reviews for ICLR 2026 submissions, finding that while LLMs can broadly distinguish accepted from rejected papers, they fail to replicate human distinctions between oral and poster presentations and exhibit provider-specific scoring biases and thematic differences in identifying weaknesses.

Original authors: Abraham Camelo-Guerrero, Jairo Diaz-Rodriguez

Published 2026-08-05
📖 4 min read☕ Coffee break read

Original authors: Abraham Camelo-Guerrero, Jairo Diaz-Rodriguez

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a world where every new scientific idea has to pass through a gatekeeper before it can be shared with the rest of the world. For decades, this gatekeeper has been a human expert, a peer reviewer, who reads a paper, ponders its strengths and weaknesses, and gives it a score. It's a bit like a talent show judge, but instead of singing, they are judging complex math and computer code. Recently, a new type of judge has entered the arena: Artificial Intelligence, specifically Large Language Models (LLMs). These are super-smart computer programs trained on vast amounts of text that can read a paper and write a review just like a human. But here is the big question: Do these AI judges actually think like human judges, or are they just mimicking the words? If an AI says a paper is "good," does it mean the same thing as when a human says it? This is the mystery scientists are trying to solve, because if we start using AI to decide which scientific papers get published, we need to know if they are playing by the same rules as the humans they are replacing.

This paper is like a giant, controlled experiment where the researchers set up a "reviewing showdown." They took 300 real scientific papers from a major computer science conference (ICLR 2026) and asked three different AI models—OpenAI's GPT-5.4, Google's Gemini 3.1 Pro, and Anthropic's Claude Opus 4.6—to review them. To make it a fair fight, the researchers matched the papers up so that for every "oral" presentation (the top-tier, spotlight papers), there was a similar "poster" paper (good, but not top-tier) and a "rejected" paper. They then compared the AI reviews against the actual human reviews and the final decisions made by the conference.

The results were a mix of "not bad" and "very different." First, the good news: all three AI models were smart enough to tell the difference between a paper that was accepted and one that was rejected. They could see the forest. However, they completely failed to see the trees. While human reviewers gave distinct scores to the "oral" papers versus the "poster" papers (recognizing that oral papers were slightly better), the AI models treated them exactly the same. They couldn't tell the difference between a "good" paper and a "great" paper.

Furthermore, the AI models didn't just score differently; they had their own unique personalities. Google's Gemini was the "nice guy" of the group, giving systematically higher scores to everyone, even handing out high marks to papers that humans had rejected. OpenAI and Claude were more like the "tough critics," especially when it came to the top-tier oral papers, where they were actually harsher than the humans.

Finally, the researchers looked at what the reviewers were complaining about. Humans and AI agreed on the big picture, like "this experiment is weak" or "the writing is unclear." But when it came to the details, they had different priorities. The AI models were obsessed with missing comparisons to other methods (baselines), while human reviewers were much more worried about how much computer power and money the research would cost.

In short, the paper suggests that while AI can be a helpful second pair of eyes to catch obvious mistakes, it isn't ready to replace human judges. It can tell you if a paper is a "pass" or a "fail," but it struggles to understand the subtle nuances that make a paper truly excellent, and it cares about different things than the humans do. Treating an AI score as exactly the same as a human score would be like trusting a robot to judge a cooking contest because it can tell the difference between a burnt cookie and a fresh one, even though it has no idea what "gourmet" tastes like.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →