← Latest papers
💻 computer science

Position on LLM-Assisted Peer Review: Addressing Reviewer Gap through Mentoring and Feedback

This position paper argues against fully automated AI reviews and advocates for a human-centered paradigm where LLMs serve as mentoring and feedback tools to cultivate reviewer expertise and address the sustainability crisis in scholarly peer review.

Original authors: JungMin Yun, JuneHyoung Kwon, MiHyeon Kim, YoungBin Kim

Published 2026-01-15
📖 5 min read🧠 Deep dive

Original authors: JungMin Yun, JuneHyoung Kwon, MiHyeon Kim, YoungBin Kim

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a massive, high-stakes talent show where thousands of new acts (AI research papers) are trying to get on stage every year. The problem? There aren't enough judges (reviewers) to watch them all, and the ones we do have are exhausted, overwhelmed, and sometimes not fully trained to give fair, helpful feedback. This is the "Reviewer Gap."

The paper argues that instead of letting robots (AI) take over the judging entirely—which would be risky and often inaccurate—we should use AI as a personal coach to help human judges get better at their jobs.

Here is the simple breakdown of their proposal:

1. The Problem: Too Many Acts, Too Few Judges

The world of AI is growing faster than a weed in spring. Conferences are getting flooded with submissions.

  • The Volume Gap: There are simply too many papers. Reviewers are tired, leading to shallow, rushed, or "check-the-box" reviews.
  • The Quality Gap: Because there are so many papers, conferences are asking junior researchers (who haven't reviewed much before) to judge. Without training, they might miss the point or be unfair, creating a cycle of bad reviews.

2. The Old (Flawed) Idea: The Robot Judge

Some people suggested letting AI write the reviews automatically. The authors say no.

  • Why? AI can "hallucinate" (make things up), misunderstand complex science, or just summarize the paper without offering real insight. It's like asking a calculator to write a poem; it might get the words right, but it misses the soul and the nuance.
  • The Risk: If we let AI write the reviews, we lose the human expert judgment that is essential for science.

3. The New Idea: The AI Coach

Instead of replacing the judge, the paper proposes using AI to train and support the human judge. Think of it like a sports team using a video analyst. The coach doesn't play the game; they help the player see their mistakes and improve.

The system has two main parts:

Part A: The Training Camp (Mentoring System)

Before a reviewer even sees a real paper, they can practice in a "safe zone" with an AI coach.

  • Guided Recognition: The AI shows the reviewer examples of good and bad reviews. It asks, "Is this review fair?" or "Is this clear?" to help the reviewer learn the rules.
  • Refinement Practice: The AI gives the reviewer a draft review that has mistakes (like being too vague or mean). The reviewer has to fix it, and the AI gives hints like, "You said the experiment was bad, but you didn't say why. Can you point to the specific data?"
  • Full Simulation: The reviewer writes a full review, and the AI gives a detailed report card on how well they followed the rules.
  • The Goal: This isn't a mandatory test. It's a voluntary way to get a "Coach's Certification," signaling that the reviewer has taken the time to learn how to be a great judge.

Part B: The Safety Net (Feedback System)

Even experts make mistakes when they are tired or rushed. This is the second part of the system.

  • The Check-Up: After a reviewer writes a draft, the AI scans it before it is sent to the author.
  • The Evidence: If the AI sees a vague comment like "The data is weak," it checks the paper and says to the reviewer: "Hey, you mentioned the data is weak, but the paper actually used the ImageNet dataset. If you think they needed more data, can you specify which dataset they should have used?"
  • The Choice: The AI doesn't force the change. It just whispers, "Here is a suggestion to make your point clearer." The human reviewer decides whether to take the advice. This ensures the human stays in charge.

4. The Golden Rules (The Rubric)

To make sure the AI coach knows what a "good review" looks like, the paper defines five simple rules:

  1. Fidelity: Did you tell the truth about what the paper actually said? (No making things up).
  2. Clarity: Is your writing easy to understand? (No confusing jargon).
  3. Fairness: Are you judging the work, not the person? (No bias against the author's school or background).
  4. Proportionality: Is your criticism fair? (Don't tear a paper apart because of one tiny typo).
  5. Constructiveness: Did you offer a way to fix the problems? (Don't just say "it's bad," say "try this instead").

5. The Big Picture: A Virtuous Cycle

The authors believe this creates a positive loop:

  • Better Training \rightarrow Better Reviews \rightarrow Better Research \rightarrow Better AI \rightarrow Even Better Tools for Reviewing.

They argue that by treating AI as a mentor rather than a replacement, we can solve the shortage of good reviews without losing the human touch. It's about helping humans become better judges, not letting machines take the gavel.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →