← Latest papers
💬 NLP

Detecting AI-Generated Content in Academic Peer Reviews

This study analyzes peer reviews from ICLR and Nature Communications to reveal a sharp increase in AI-generated content starting in 2022, reaching approximately 20% and 12% respectively by 2025, thereby highlighting the growing prevalence of AI assistance in scholarly evaluation.

Original authors: Siyuan Shen, Kai Wang

Published 2026-03-24
📖 4 min read☕ Coffee break read

Original authors: Siyuan Shen, Kai Wang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the world of academic publishing as a massive, high-stakes Tournament of Ideas. In this tournament, scientists submit their research papers, and a group of expert judges (the peer reviewers) reads them to decide if they are good enough to be published. For decades, this was a strictly human game: humans wrote the papers, and humans wrote the critiques.

But recently, a new player has entered the arena: Artificial Intelligence (AI).

This paper is like a detective story where the authors try to figure out: How many of these expert judges are secretly letting a robot write their critiques for them? And more importantly, when did this start happening?

Here is the breakdown of their investigation, explained with some everyday analogies:

1. The Detective's Tool: The "AI Sniffer"

The researchers couldn't just ask the reviewers, "Did you use AI?" because reviewers are anonymous and might lie. Instead, they built a digital metal detector.

  • How they built it: They fed the detector a pile of reviews from 2021. They told the detector, "Here are the real human reviews, and here are some fake reviews made by AI. Learn the difference."
  • The Training: The detector learned to spot the "fingerprint" of AI writing. It's like teaching a dog to smell a specific type of cheese; once trained, the dog can sniff out that cheese even if it's hidden in a different room.

2. The Investigation: Scanning the Past

Once the detector was ready, the researchers used it to scan reviews from 2022, 2023, 2024, and 2025. They looked at two different "arenas":

  • ICLR: A fast-paced, annual conference for computer science (like a sprint).
  • Nature Communications: A prestigious, continuous journal for all sciences (like a marathon).

3. The Findings: The "Silent" Years vs. The "Explosion"

The results were like watching a quiet pond suddenly get a massive splash.

  • 2022 & 2023 (The Quiet Years): The detector barely found anything. It was like walking through a forest and finding zero plastic bottles. Almost all reviews looked like they were written by humans. This suggests that back then, using AI to write reviews was rare or non-existent.
  • 2024 (The Turning Point): Suddenly, the detector started beeping. In the second half of 2024, the number of AI-sounding reviews started climbing. It's as if someone opened a floodgate.
  • 2025 (The Flood): By 2025, the situation had changed dramatically.
    • At the ICLR conference, nearly 20% of the reviews (1 in 5) were flagged as AI-generated.
    • At Nature Communications, about 12% were flagged.

4. The "Why" and the "When"

The paper noticed something interesting about the timing. The biggest jump happened between the third and fourth quarters of 2024.

Think of this like a new video game console being released. In May 2024, OpenAI released a super-smart version of their AI (ChatGPT-4o) that was free for everyone to use. The data suggests that as soon as this "super-tool" became available, reviewers started using it more and more to draft their feedback.

5. The Big Picture: What Does This Mean?

The authors warn us that this isn't just a glitch; it's a shift in behavior.

  • The "Ghost Writer" Problem: Imagine if a referee in a sports game let a robot write the penalty calls. The game might still look okay, but the integrity is compromised. If AI writes the reviews, are the scientists actually reading the papers, or are they just hitting "generate"?
  • The "Hybrid" Reality: The paper admits their detector is looking for reviews that are mostly AI. In reality, many reviewers might be doing a "human-AI dance"—writing the main ideas themselves but using AI to polish the grammar or make it sound fancier. Their detector might miss these "hybrid" cases, meaning the real number of AI usage could be even higher.

The Takeaway

This study is a wake-up call. It shows that the academic world is moving from a "human-only" zone to a "human-and-robot" zone very quickly.

Just as schools had to update their rules when calculators and then the internet arrived, the world of science needs to figure out new rules for AI. If 1 in 5 reviews is written by a machine, we need to ask: Are we still judging the quality of science, or are we just judging how well the AI can mimic a human?

The authors conclude that while AI might help efficiency, we need to be careful not to let the "ghosts in the machine" take over the most important job of all: deciding what new knowledge gets to shape our future.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →