← Latest papers
💬 NLP

Is Peer Review Really in Decline? Analyzing Review Quality across Venues and Time

Challenging the popular narrative of declining standards, this paper introduces a new framework for quantifying review quality and demonstrates through a cross-temporal analysis of major AI and linguistics conferences that median review quality has remained stable over time.

Original authors: Ilia Kuznetsov, Rohan Nayak, Alla Rozovskaya, Iryna Gurevych

Published 2026-01-22
📖 5 min read🧠 Deep dive

Original authors: Ilia Kuznetsov, Rohan Nayak, Alla Rozovskaya, Iryna Gurevych

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Question: Are Reviews Getting Worse?

Imagine you are running a massive, global talent show. Every year, thousands of singers (researchers) submit their songs (papers) to be judged by a panel of experts (reviewers). Recently, there has been a lot of gossip in the audience: "The judges are getting lazy! They aren't listening closely anymore. The quality of the feedback is dropping because there are too many singers and not enough judges."

This paper asks a simple question: Is this gossip true? Do the reviews actually get worse over time, or is that just a feeling?

The Problem: You Can't Compare Apples to Oranges

To answer this, the researchers tried to measure "review quality." But they hit a snag. Imagine trying to compare the quality of a sandwich from 2018 to a sandwich from 2025.

  • In 2018, the sandwich might have been served on a napkin with no instructions.
  • In 2025, it comes in a fancy box with a checklist of ingredients.
  • Some judges write long paragraphs; others write bullet points.
  • Some judges write in a specific format; others write freely.

Because the "forms" and styles changed so much over the years and across different conferences (like ICLR, NeurIPS, and ACL), it was impossible to fairly compare them. It was like trying to measure the height of people using a ruler for one group and a tape measure for another.

The Solution: The "Review Translator"

To fix this, the authors built a universal translator for reviews.

  1. Flattening: They took all the different forms and flattened them into one long, continuous block of text.
  2. Itemization: They used a smart AI (a Large Language Model) to act like a professional editor. This AI took the messy text and reorganized it into standard sections: Summary, Strengths, Weaknesses, and Other.

Think of this as taking every sandwich, regardless of how it was served, and slicing it into standard pieces so you can taste the bread, the meat, and the cheese separately. This allowed them to compare reviews from 2018 directly with reviews from 2025.

How They Measured "Quality"

Once the reviews were standardized, the team measured them using three main "taste tests":

  1. Substantiveness (Is there enough meat on the bone?):

    • Lightweight: How long is the review? (Counting characters).
    • Deep: How many distinct points or "items" did the reviewer make?
    • Analogy: A short review is like a tiny cracker; a long one is a full meal. They counted how many "bites" of information were in the review.
  2. Actionability (Can the singer actually fix their song?):

    • Did the reviewer give specific instructions? (e.g., "Change Figure 3" vs. "This is bad").
    • Analogy: A good review is like a coach saying, "Your left foot is too far forward." A bad review is just shouting, "You're off!" The authors measured how many specific "fix-it" requests were in the text.
  3. Grounding (Did they look at the actual paper?):

    • Did the reviewer point to specific parts of the paper? (e.g., "See page 4, line 10").
    • Analogy: A grounded review is like a detective pointing to evidence on the table. An ungrounded review is just guessing without looking at the clues.

The Results: The Gossip Was Wrong

After analyzing thousands of reviews from top computer science conferences over several years, the data told a surprising story:

The median quality of reviews did NOT decline.

Despite the fear that reviewers are burning out or that AI is making them lazy, the "average" review today is just as helpful, detailed, and grounded as the reviews from a few years ago. The "gossip" of decline wasn't supported by the numbers.

Why Does It Feel Like Quality is Dropping?

If the quality is the same, why does everyone feel like it's getting worse? The authors offer a few theories (hypotheses):

  • The "More People, More Noise" Theory: Even if the average quality stays the same, as the number of submissions grows, the total number of bad reviews increases. If you have 100 reviews and 5 are bad, that's 5%. If you have 10,000 reviews and 500 are bad, that's still 5%, but suddenly 500 people are angry. The volume of bad experiences makes it feel like everything is broken.
  • The "Worst Got Worse" Theory: Maybe the average is fine, but the absolute worst reviews are becoming even worse, and those are the ones people talk about on social media.
  • The "We Just Don't Know" Theory: Maybe the quality did drop in ways the authors couldn't measure (like the "soul" of the review), but their current tools couldn't see it.

The Takeaway

The paper concludes that we shouldn't panic and assume the system is collapsing. The "average" review is holding its ground. However, because the number of papers is exploding, we need to work harder to ensure that everyone gets a good review, not just the average one.

The authors also built a toolkit (code and data) that other researchers can use to keep checking this "health" of the review system in the future, ensuring we have facts, not just feelings, to guide us.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →