← Latest papers
💬 NLP

AutoSupervision: Closing the Feedback Loop in Scientific Workflows with Grounded Revision Verification

This paper introduces AutoSupervision, a framework leveraging 56,000 Nature Communications review records to evaluate whether manuscript revisions genuinely address reviewer concerns through grounded evidence, revealing that while large language models can effectively characterize concerns, evidence-based verification remains a significant bottleneck.

Original authors: Haobo Li, Eunseo Jung, Wenxiao Zhao, Feng Liu, Jiong Wang, Kaiyi Xu, Zijie Guo, Zixin Chen, Ben Fei, Fenghua Ling, Lei Bai

Published 2026-07-31
📖 3 min read☕ Coffee break read

Original authors: Haobo Li, Eunseo Jung, Wenxiao Zhao, Feng Liu, Jiong Wang, Kaiyi Xu, Zijie Guo, Zixin Chen, Ben Fei, Fenghua Ling, Lei Bai

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine science as a giant, never-ending game of "telephone" played by the smartest people on Earth, but with a twist: instead of whispering, they write down their ideas, and instead of just passing them along, they have to prove they actually changed the message when someone points out a mistake. This is the world of peer review, the rigorous checkpoint where scientists submit their discoveries, and other experts act as strict editors, pointing out holes in the logic, missing experiments, or confusing explanations. For a long time, we've had computers that are pretty good at generating these scientific stories or even critiquing them like a harsh editor. But there's been a missing piece in the puzzle: can a computer actually look at the final draft, compare it to the editor's original complaints, and say, "Yes, the author actually fixed the problem, and here is the exact sentence where they did it"? This is the tricky part of the "feedback loop"—verifying that the changes made are real, meaningful, and backed by evidence, not just empty promises.

Enter AutoSupervision, a new tool built by researchers to test exactly this ability. Think of it as a super-smart "fact-checker" for scientific papers. The team created a massive training ground using 56,000 real articles and their review records from Nature Communications. They set up a challenge where an AI has to read a reviewer's complaint, the author's reply, and the new version of the paper, then decide: Did the author actually fix the issue? If so, where in the text is the proof? The results are a bit of a mixed bag. The AI models are fantastic at understanding what the reviewer was worried about—almost like a perfect student who can summarize the teacher's notes. However, when it comes to the hard work of verifying if the changes were actually made and finding the specific evidence in the text, the AI stumbles. The best-performing model managed to get a score of only 0.501 on this verification task, while its ability to just understand the problem was much higher at 0.754. Essentially, the computers are great at listening to the problem but still struggle to prove the solution is real.

The researchers found that while modern AI can almost perfectly identify the concerns (a "coverage" score near 1.000), it often fails to connect the dots between the author's claim ("We fixed it!") and the actual text in the revised manuscript. It's like a student who confidently tells the teacher, "I did my homework," but when asked to show the work, they can't point to the right page. The study suggests that while AI is ready to help generate feedback, it isn't quite ready to be the final judge of whether that feedback was truly acted upon. The team also discovered that giving the AI a little help—like breaking the task into smaller steps or training it specifically on these kinds of problems—does improve its performance, but the "grounding" problem (finding the exact evidence) remains the biggest hurdle. In short, we have AI that can read the room, but we still need to teach it how to check the receipts.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →