Policies Permitting LLM Use for Polishing Peer Reviews Are Currently Not Enforceable
This paper demonstrates that current AI-text detectors are not enforceable for identifying LLM use in peer reviews because they frequently misclassify human-written reviews that have been polished by AI as fully AI-generated, leading to false accusations of academic misconduct and unreliable estimates of policy violations.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a high-stakes game of Peer Review as a massive, global Taste-Testing Competition for scientific papers.
In this competition, expert judges (reviewers) read a new recipe (the paper) and write a critique: "This dish is delicious," or "This sauce is too salty." Recently, a new tool has entered the kitchen: AI Chefs (Large Language Models).
Some competition organizers (conferences and journals) have made a new rule:
"You can use the AI Chef only to fix your grammar, smooth out your sentences, and make your writing flow better. But you must write the actual opinion and arguments yourself."
This is the "Polishing-Only" Policy. It's like saying, "You can use a food processor to chop your onions, but you can't let the robot cook the whole meal."
The big question this paper asks is: "Can we actually catch people who break this rule?"
The authors built a giant Detective Lab to test if current AI detectors can tell the difference between:
- The Human Chef: Wrote the whole review.
- The AI Chef: Wrote the whole review (a clear violation).
- The Hybrid Chef: Wrote the review, but used the AI to polish the grammar (allowed).
Here is what they found, using some simple analogies:
1. The "False Alarm" Problem (The Fire Alarm that Goes Off for Toast)
The researchers tested five different "AI Detectors" (like advanced smoke alarms).
- The Good News: If a reviewer let the AI write the entire review, the detectors were great at catching it. They screamed "FIRE!" almost every time.
- The Bad News: When a reviewer wrote the review themselves and just asked the AI to "fix the grammar" (the allowed polishing), the detectors still screamed "FIRE!" about 3% to 6% of the time.
The Analogy: Imagine a security guard at a concert. If someone sneaks in a full band (AI-generated review), the guard catches them. But if a fan brings a small, allowed guitar pick (AI-polished review), the guard sometimes mistakes it for a full band and kicks them out.
- The Consequence: If conferences enforced the rule using these detectors, they would falsely accuse innocent reviewers of cheating. In a big conference with 75,000 reviews, a small error rate means thousands of people could be wrongly accused of academic misconduct.
2. The "Chameleon" Effect
The paper also looked at "Humanized" reviews. This is when someone tries to trick the detector by running their AI-written review through a "humanizer" tool (like putting a disguise on the AI).
- The Result: The detectors got confused. They stopped catching the bad AI reviews, but they also started flagging more innocent, polished reviews as fake. It was a mess.
3. Trying to Cheat the System with "Context"
The authors thought, "Wait! We have an advantage! We have the original paper the review is about. Maybe we can use that to catch the AI."
- The Idea: If the AI wrote the review, it might sound too similar to the paper's style, or it might miss a subtle nuance that a human would catch.
- The Reality: They tried using the paper's text to help the detectors. It helped a little bit, but not enough. The detectors still couldn't reliably tell the difference between a "Human + Polish" review and a "Full AI" review. It's like trying to tell if a song was written by a human or a robot just by looking at the sheet music; sometimes they look identical.
4. The "Magic 21%" Claim
The paper mentions a recent news story where a company claimed 21% of reviews at a major conference were written entirely by AI.
- The Paper's Take: The authors say, "Hold on a second." Because the detectors are so bad at distinguishing between "Full AI" and "AI-Polished Human," that 21% number is likely wrong. It probably includes many reviewers who just used AI to fix their grammar. The "AI invasion" might be much smaller than reported, or the detectors are just too noisy to give a real number.
The Big Conclusion
The paper concludes that we currently cannot enforce the "Polishing-Only" policy.
It's like trying to enforce a rule that says "You can use a calculator for addition, but not for multiplication," but your calculator is so broken it sometimes thinks addition is multiplication.
- If you ban AI completely: You might miss the people who used it to write the whole thing (because the detectors aren't perfect).
- If you allow polishing: You can't catch the people who are actually cheating, because the detectors will also catch the innocent people who just polished their work.
The Takeaway for Everyone:
Until we build much smarter "detective tools," we have to be very careful about accusing people of cheating based on AI detectors. The technology isn't ready for the courtroom of academic peer review yet. The best advice for reviewers? If you use AI to polish your work, be very clear in your instructions to the AI ("Don't add new ideas, just fix the grammar") and double-check the result yourself!
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.