Evidence-processing errors and their correction in an LLM-assisted systematic review: a retrospective methodological case study
This retrospective methodological case study investigates how documented evidence-processing errors affected an LLM-assisted systematic review, demonstrating that an auditable workflow successfully linked error detection and correction to released outputs while providing operational rules for future review teams.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
Imagine trying to solve a massive puzzle where the pieces are scattered across thousands of different books, and some of those books describe the exact same group of people. This is the daily reality of a systematic review, a rigorous scientific method used to gather all available evidence on a specific medical question. Researchers do this to find the true answer hidden among conflicting studies, but the process is incredibly fragile. A single mistake—like misreading a date, accidentally counting the same patient twice, or misunderstanding whether a result is good or bad—can twist the final conclusion. Now, imagine adding artificial intelligence to help sort through this mountain of information. While these computer programs can read and organize data at incredible speed, they are not perfect. They can make subtle errors that are hard to spot, and if those errors slip through, the entire scientific conclusion could be wrong. The crucial question for modern science is not just whether the computer can do the work, but how we can catch its mistakes and ensure the final answer is trustworthy.
A team of researchers in China decided to answer this question by looking backward at a real-world project they had just finished. They had used a sophisticated system that combined large language models—advanced AI tools capable of understanding human language—with strict human oversight to conduct a review on how social connections before surgery affect recovery afterward. Instead of just presenting their final medical findings, they turned the spotlight on the process itself. They treated their own completed project as a laboratory to investigate exactly where things went wrong, how those errors were found, and how they were fixed before the results were released to the public. Their goal was not to prove that their specific review was perfect, but to create a transparent map of the errors that occurred and the safeguards that stopped them from ruining the science.
The researchers examined the entire history of their project, from the initial search for studies to the final publication. They discovered fifty distinct root-cause events, which are the fundamental reasons why a mistake happened. These were not just minor typos; some of these errors had already changed the combined statistical results of the review before they were caught and corrected. For instance, in one case, the system had included data from patients starting their follow-up at the wrong time, which skewed the survival statistics. In another, the AI failed to realize that two different reports described the same group of patients, leading to those patients being counted twice, which artificially inflated the weight of that evidence. The team also found errors where the computer misinterpreted the direction of a result, thinking a negative outcome was a positive one, or where it selected the wrong type of data to analyze.
What makes this study particularly valuable is how the team handled these errors. They did not simply delete the mistakes and pretend they never happened. Instead, they built a system that acted like a detailed audit trail. Every time a mistake was found, they recorded exactly what went wrong, what the consequence was, and how they fixed it. They traced four specific pathways of failure to show how a small error in understanding a source document could ripple through the entire system, changing which studies were included, how much weight they carried, and even the final confidence rating of the result. For example, when they realized an overlap between studies had been missed, they corrected the link, which changed the number of studies included in the calculation and slightly adjusted the final odds ratio. In another instance, a correction regarding the risk of bias in a study changed the assessment of the evidence's quality, even though the final grade remained at the lowest level.
The researchers found that most of these errors were caught by a combination of automated checks and human review. The system was designed so that if a rule was broken—such as counting a patient twice or using the wrong time frame—the process would stop and flag the issue. Humans then stepped in to make the final judgment. In the end, they resolved forty-six of the fifty errors, while four were acknowledged as disclosed non-critical limitations that did not change the main conclusions. The final review included 454 reports representing 445 unique studies, and despite the fifty errors that occurred during the process, the team was able to correct the resolved ones before releasing the final scientific output. They verified their work by running the entire process again with a different set of code and by having two independent reviewers check the same source documents, confirming that the final numbers were consistent and the files matched perfectly.
This work does not claim that artificial intelligence is ready to replace human scientists or that these tools are flawless. In fact, the study explicitly shows that relying solely on AI without a clear, auditable process would have led to incorrect results. The researchers emphasize that their findings are specific to this one project and cannot be used to calculate a general error rate for all AI systems. However, they provide a concrete set of rules and examples that other research teams can use to build their own safety nets. By documenting exactly how errors happen and how they can be traced back to their source, the team has shown that it is possible to use powerful AI tools in high-stakes science, provided there is a rigorous system in place to catch mistakes, correct them, and prove that the final answer is reliable. The study serves as a practical guide for how to build trust in a future where computers help us read the world's medical literature.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.