Local AI pre-screening for human triple-blind peer review in health sciences
This paper proposes a transparent, triple-blind framework for health sciences peer review that utilizes locally-hosted, open-weight LLMs for multi-stage pre-screening to mitigate risks associated with undisclosed AI use and reviewer shortages, while ensuring human experts retain final decision-making authority.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Science relies on a quiet, essential ritual: before a new discovery is shared with the world, other experts must read it carefully to ensure it is sound. This process, known as peer review, is how the scientific community decides which work is worthy of publication. For decades, this system has functioned by matching a submitted paper with a handful of human experts who read it in secret, offering their judgment on whether the methods are correct and the conclusions hold up. But the system is currently buckling under its own weight. In recent years, the number of papers submitted to major conferences has exploded, reaching numbers like 21,575 for a single event, while the number of qualified experts willing to read them has not kept pace. The result is a bottleneck where good science waits for months, and the pressure has become so great that some researchers have quietly begun using artificial intelligence to write these reviews, often without telling anyone. This hidden shift brings new risks, such as computers inventing fake references or being tricked by authors who hide instructions inside their text to force a positive rating.
In response to this growing crisis, a researcher named Rodrigo M. Boos has proposed a new way to organize the review process that brings the use of artificial intelligence into the open while keeping human experts firmly in charge. The paper outlines a system designed specifically for health science journals, where the stakes for patient safety and research integrity are exceptionally high. The core idea is to use a team of three artificial intelligence programs to perform an initial screening of every paper before a single human reviewer ever sees it. These programs do not make the final decision to publish or reject a paper; instead, they act as a filter, checking if a manuscript is even ready for human attention. Crucially, the system is built to protect the privacy of the unpublished work. Unlike many commercial tools that send data to the internet, this proposal requires the artificial intelligence to run on the journal's own local computers, ensuring that the text of a new medical study never leaves the journal's secure infrastructure. This design directly addresses the concerns of major funding agencies that currently ban their reviewers from using outside artificial intelligence tools because they cannot guarantee where that unpublished data is stored.
The system works by guiding a submitted manuscript through five distinct stages. First, the paper is automatically cleaned to remove any hidden text or secret instructions that might try to trick the computer. Next, three independent artificial intelligence models read the anonymized paper and score it against a specific set of criteria, such as whether the study design supports the conclusions or if the statistics are sound. These models are not allowed to make a final judgment on their own; they simply provide a score and a structured explanation. If the scores are too low, the paper is sent back to the author with feedback on what needs fixing. If the scores are high enough, the paper moves to the next stage, where three human experts read it. These human reviewers are blind to the scores given by the artificial intelligence; they only know that the paper passed the initial computer check. This separation is vital to prevent the humans from simply agreeing with the computer's opinion. Finally, an editor reviews the work of both the artificial intelligence and the humans to make the final decision on publication.
A particularly novel part of this proposal is a feature that allows authors to choose how they want to be handled if the paper is not ready. In a standard review, an author receives a rejection letter from a human editor, which can be a painful and stigmatizing experience that discourages them from continuing their work. This new system offers an option where an author can agree in advance that if the three artificial intelligence models all agree the paper is not ready, the paper will be returned to them for revision without any human ever seeing it. This creates a private, "unwitnessed" rejection that spares the author the social shame of having a human judge their work before it is polished. The author of the paper acknowledges that this is a hypothesis that needs testing; it is not yet proven that this approach reduces the emotional burden of rejection, but it is a deliberate attempt to address a known psychological harm in the current system.
The paper is careful to state that this is a design proposal, not a finished product that has been tested on real medical journals yet. The author admits that the specific scores used to decide if a paper passes the computer check are working guesses that will need to be calibrated once the system is actually used. There are also known limitations: the artificial intelligence models used in this local system are slightly less powerful than the most advanced commercial models available on the internet, a trade-off made to ensure total privacy. Furthermore, the author notes that while the system is designed to catch obvious errors, it cannot yet guarantee that it will never miss a subtle flaw or be tricked by a sophisticated attack. The proposal does not claim to replace human judgment, which remains the mandatory final step, but rather to act as a structured, transparent layer that helps manage the overwhelming volume of submissions. By making the role of artificial intelligence visible and controlled, the author argues that the scientific community can use these tools to speed up the process without sacrificing the trust and integrity that peer review is meant to protect.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.