← Latest papers
🤖 AI

Position: Want Better ML Reviews? Stop Asking Nicely and Start Incentivizing with a Credit System

This position paper argues that improving machine learning peer review requires moving beyond polite guidelines to implement enforceable procedural safeguards and a credit-based incentive system, such as "OpenReview Points," to reward quality contributions and regulate submission volumes.

Original authors: Shaochen Zhong

Published 2026-08-18
📖 7 min read🧠 Deep dive

Original authors: Shaochen Zhong

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the vast and rapidly expanding world of machine learning, a field dedicated to teaching computers to learn from data, a critical bottleneck has emerged not in the technology itself, but in the human process used to validate it. Before a new discovery can be shared with the world, it must pass through peer review, a system where experts read and evaluate the work of their colleagues to decide if it is worthy of publication. For decades, this process has relied on a simple social contract: if you want your work published, you must agree to review the work of others. However, as the number of research papers has exploded into the tens of thousands for a single major meeting, this polite agreement has begun to fray. The sheer volume of submissions has overwhelmed the available experts, leading to rushed evaluations, inconsistent standards, and a growing sense of frustration among scientists who feel the system is no longer functioning as intended. The core question facing the community is no longer just how to manage the workload, but how to create a system that actually encourages high-quality, thoughtful feedback rather than just checking a box.

A recent position paper by Shaochen (Henry) Zhong of Rice University argues that the current approach of asking researchers to be more careful or to follow better guidelines is insufficient. The author suggests that the machine learning community has reached a point where polite requests and optimistic guidelines are failing to solve the crisis of submission volume and review quality. Instead, the paper proposes a radical shift toward a system that treats reviewing as a form of currency. The central idea is to replace the current "honor system" with a tangible credit economy, where researchers earn points for their contributions to the review process and can spend those points to gain specific privileges. This approach moves away from relying on goodwill and toward a structure where good behavior is rewarded and poor behavior carries a measurable cost.

The paper identifies two main drivers of the current dysfunction. First, the number of papers submitted to major conferences has grown so large that it is physically impossible for the available experts to give each one the attention it deserves. With some conferences receiving over twenty thousand submissions, the system is forced to make quick, often arbitrary decisions to keep the process moving. Second, there is a lack of consequences for bad behavior. A reviewer can submit a dismissive, low-effort, or even nonsensical critique without fear of repercussion, while a reviewer who goes the extra mile to provide detailed, constructive feedback receives no tangible recognition. The author notes that while some conferences have tried to limit submissions by capping the number of papers an author can submit, these measures often fail because they do not address the root cause: the lack of a downside for submitting unready work or the lack of a reward for doing the hard work of reviewing well.

To address these issues, the author proposes a system called "OpenReview Points." In this model, researchers would earn points by performing specific tasks, such as writing a standard review, helping with an emergency review, or being recognized as an outstanding reviewer. These points would then function as a flexible currency that could be spent across different major conferences. For example, a researcher could spend points to request an exemption from a review duty, to ask for an additional expert to weigh in on a controversial paper, or even to redeem a free registration fee for a conference. The system is designed to be granular and enforceable, allowing organizers to reward good behavior and penalize bad behavior in ways that are more nuanced than the current "all-or-nothing" approach.

The paper explicitly argues against several existing solutions that the community has tried. It suggests that simply asking reviewers to be nicer or more thorough is ineffective because it relies on motivation that the current system does not support. It also critiques the idea of strict, hard limits on the number of papers an author can submit, noting that such rules often lead to teams simply removing names from papers to fit under the cap, rather than actually reducing the number of submissions. Furthermore, the author warns that harsh penalties, such as rejecting a paper because a reviewer was low-effort, are too blunt an instrument to handle the wide range of poor reviewing behaviors. Instead, the proposed credit system offers a middle ground where penalties can be proportional, such as deducting points for a low-quality review, which serves as a deterrent without destroying a researcher's career.

A key component of the proposal is the idea of "fine-grained procedural safeguards." The author outlines four specific mechanisms to prevent the system from being abused. First, the system would restrict certain actions to specific roles, ensuring that only the lead authors of a paper can spend points to request extra help, preventing the burden from being shifted to less accountable team members. Second, there would be upper limits on how often certain actions can be taken, preventing anyone from "farming" points by submitting a massive number of low-effort reviews. Third, the system would use dynamic pricing, meaning that the more a researcher uses a particular privilege, the more it costs in points, ensuring that scarce resources are reserved for the most critical needs. Finally, the system would rely on peer voting to judge the quality of reviews, allowing the community to collectively identify and penalize bad actors while rewarding those who provide exceptional feedback.

The author acknowledges that this system is not without its challenges and potential downsides. There is a concern that a credit system could favor researchers who already have more time and resources, allowing them to "buy" their way out of responsibilities. However, the paper argues that because points are earned through labor rather than purchased with money, the system rewards effort rather than wealth. There is also the risk that the system could be manipulated or that the points could lose value over time, similar to inflation in a real economy. To counter this, the author suggests that points should expire after a certain period and that the system should be rolled out gradually, starting with existing perks like free conference registration before introducing more complex features.

The paper concludes by emphasizing that the goal is not to create a perfect system, but to make the current one more sustainable and accountable. The author suggests that the first conferences to adopt this system should start small, using existing rewards to test the waters before expanding the scope. By tracking key metrics and sharing data across different conferences, the community can learn what works and what does not, eventually building a shared framework that benefits everyone. The ultimate vision is a peer review process where good behavior is recognized and rewarded, where researchers have a stake in the quality of the system, and where the frustration of the current era is replaced by a more functional and fair exchange of ideas. This proposal represents a shift from relying on the hope that researchers will do the right thing to creating a structure where doing the right thing is the most logical and beneficial choice.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →