Impact of large language models on peer review opinions from a fine-grained perspective: Evidence from top conference proceedings in AI
This study reveals that the emergence of Large Language Models has altered peer review in top AI conferences by making reports longer, more fluent, and standardized, while simultaneously shifting focus toward surface-level clarity and summaries at the expense of deeper critical evaluations regarding originality and replicability.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the world of academic research as a massive, high-stakes talent show. Thousands of scientists submit their "acts" (research papers) to top-tier judges (peer reviewers) at prestigious venues like ICLR and NeurIPS. The judges' job is to read these acts, critique them, and decide who gets a spot on the main stage (publication).
For decades, this process has been run entirely by humans. But recently, a new, incredibly smart AI assistant (Large Language Models, or LLMs like ChatGPT) has entered the judges' booth.
This paper asks: What happens when judges start using this AI assistant to write their critiques? Do they become better judges, or do they just sound different?
Here is the breakdown of the study, explained through simple analogies:
1. The Setup: A Flood of Submissions
The "talent show" has become so popular that there are too many acts and not enough judges. The judges are overwhelmed, tired, and sometimes biased. To cope, they started using AI tools to help them write their feedback reports. The researchers wanted to see if this changed the quality of the judging, not just the speed.
2. The Investigation: Looking Under the Hood
The researchers didn't just look at the final scores; they looked at the words the judges wrote. They used a "microscope" to examine the reviews in three ways:
- The Length: Are the reviews getting longer or shorter?
- The Vocabulary: Are the words getting fancier or simpler?
- The Focus: What are the judges actually talking about? (e.g., Is the paper original? Is the math sound? Is the summary clear?)
They also used a special "lie detector" (a statistical model) to guess which reviews were likely helped by AI and which were written purely by humans.
3. The Findings: The "AI-Flavored" Review
📝 The "Longer and Smoother" Effect
The Analogy: Imagine a human judge writing a critique. It might be a bit choppy, with short sentences and some rough edges. Now, imagine an AI polishing that critique. It becomes a long, flowing river of text.
- What happened: Reviews that used AI became longer and smoother. They sounded more professional and fluent.
- The Catch: While they sounded better, they sometimes used simpler vocabulary. It's like the AI smoothed out the jagged rocks, but also removed some of the unique, complex stones that made the river interesting.
🎯 The "Summary vs. Deep Dive" Shift
The Analogy: Think of a movie review. A human critic might spend 10 minutes analyzing the plot twists and character depth (the "Originality"). An AI-assisted critic might spend 10 minutes summarizing the plot (the "Summary") because the AI is really good at summarizing.
- What happened: AI-assisted reviews talked much more about summaries and surface-level clarity.
- What was lost: They talked less about deep, critical things like "Is this idea truly new?" (Originality) or "Can others repeat this experiment?" (Replicability).
- The Metaphor: The AI helped the judges write a better book report, but they became less likely to write a deep literary analysis.
🧠 The "Confidence" Factor
The Analogy: Imagine a group of judges. Some are Experts (Confidence Score 5) who know the field inside out. Others are Novices (Confidence Score 1) who are a bit unsure.
- The Experts: They started writing shorter, more concise reviews. They didn't need the AI to tell them what was obvious.
- The Novices: They started writing much longer reviews. The AI acted as a "training wheel," helping them understand the paper and write more detailed feedback than they could have on their own.
- The Result: The gap between the experts and the novices narrowed. The AI helped the less confident judges sound more like the experts.
4. The Verdict: Did the Scores Change?
The researchers asked: "Does using AI make the judges give higher or lower scores?"
- The Answer: Surprisingly, no major change.
- The Insight: The AI helped the judges write the report better (better grammar, better structure), but it didn't seem to change their judgment of the paper's value. The scores remained mostly the same. The AI was a scribe, not a decision-maker.
5. The Big Picture: A Double-Edged Sword
The Good News:
- Reviews are clearer and easier to read.
- Less confident judges can provide better feedback with AI help.
- The process is more efficient.
The Bad News:
- We might be losing the "spark" of deep, critical thinking. If everyone uses the same AI to write reviews, the reviews might start to sound the same (homogenized).
- We might stop asking the hard questions about whether an idea is truly new or reproducible, focusing instead on whether the text is clear.
The Final Takeaway
The study concludes that AI is like a powerful spell-checker and editor for academic reviews. It makes the text look polished and professional, but it doesn't necessarily make the judge smarter or more critical.
The danger isn't that AI will take over the judging; the danger is that we might get so used to the "smooth, AI-polished" reviews that we forget to look for the deep, messy, human insights that actually drive science forward. The researchers suggest we should embrace these tools but keep a close eye to ensure we don't lose the "soul" of critical peer review.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.