CHARMS and PROBAST+AI: an updated template for Data Extraction and Risk of Bias Assessment in systematic reviews of prediction models
This paper presents an updated, open-access Excel template that integrates the CHARMS checklist with the new PROBAST+AI framework to provide a standardized, automated tool for data extraction and risk of bias assessment in systematic reviews of both traditional and AI-driven clinical prediction models.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
In the modern medical landscape, doctors increasingly rely on mathematical tools to predict a patient's future health. These tools, known as clinical prediction models, analyze a person's symptoms, test results, and history to estimate the likelihood of a disease or the success of a treatment. For years, researchers have used standard statistical methods to build these models, but the field is rapidly shifting. Today, powerful computer systems capable of learning from vast amounts of data—often called artificial intelligence—are being used to create even more complex predictions. While these new methods hold great promise, they introduce new ways for errors to creep in. A model might appear to work perfectly on the data it was trained with but fail miserably when applied to real patients, a problem often caused by the computer accidentally "seeing the answers before the test begins" by seeing the answers before the test begins. To ensure these tools are safe and reliable, scientists must rigorously check every step of how they are built and tested. This process is known as a systematic review, where experts gather all available studies on a topic and evaluate their quality using strict checklists.
The challenge for these reviewers has been that the rules for checking old statistical models did not fully account for the quirks of modern artificial intelligence. A team of researchers from the University of Glasgow and the Arab American University has addressed this gap by creating a new, updated digital tool designed to streamline this critical work. They have taken an existing spreadsheet template, which was already used to organize data and check for bias, and completely overhauled it to match the latest guidelines for evaluating artificial intelligence. The result is a sophisticated yet user-friendly Excel file that allows researchers to separate the assessment of how a model was built from the assessment of how well it was tested. This distinction is vital because a model can be built with high-quality data but then tested in an inadequate way, or vice versa. By keeping these two evaluations distinct, the new tool ensures that reviewers do not miss subtle flaws that could render a medical prediction tool dangerous or useless.
The core innovation in this new template is its ability to handle the specific pitfalls of machine learning without requiring the reviewer to be a computer scientist. The spreadsheet now includes specific questions about whether the computer was allowed to peek at the test data during the training phase, a mistake known as data leakage. It also asks whether the researchers properly adjusted the model if the data they used had an uneven mix of healthy and sick people, a common issue that can skew results. Furthermore, the tool checks if the entire process of building the model was repeated correctly when the researchers tested it, ensuring that the final numbers reflect reality rather than an accidental fluke. These questions are woven directly into the spreadsheet so that as a reviewer enters information about a study, the tool automatically highlights where the study might have gone wrong.
What makes this tool particularly powerful is how it manages the flow of information. Instead of forcing a reviewer to flip back and forth between different documents to find the same piece of information, the spreadsheet links everything together. When a reviewer enters details about a study's participants or the data they used in one section, that information automatically appears in the sections where the quality of the model is being judged. This reduces the chance of human error and saves significant time. The tool is designed to handle up to thirty different prediction models in a single file, making it suitable for large-scale reviews. As the reviewer works through the checklist, the spreadsheet automatically generates summary tables and charts that show exactly how many models passed the quality checks and how many failed. These visual summaries update instantly, giving the research team a clear, real-time picture of the evidence they are reviewing.
The researchers emphasize that this tool is not limited to artificial intelligence; it works just as well for traditional statistical models. This universality is crucial because many medical reviews now include a mix of both old and new methods. By providing a single, standardized framework, the tool ensures that a model built with a simple equation is judged by the same rigorous standards as one built with a complex neural network. The spreadsheet is free for anyone to download and use, and it operates without the need for complex computer programming code, making it accessible to researchers who may not have advanced technical skills. The authors note that while the tool automates the heavy lifting of data organization and calculation, it does not replace the need for human judgment. The reviewer must still interpret the answers and decide what they mean for the safety and reliability of the medical tool.
In the end, this updated template represents a significant step forward in how medical evidence is gathered and verified. It acknowledges that the landscape of medical prediction is changing and provides the necessary infrastructure to keep pace. By separating the evaluation of model creation from model testing and by adding specific checks for the unique risks of artificial intelligence, the tool helps ensure that the prediction models guiding patient care are built on solid ground. The researchers hope that by making the process of checking these models easier and more consistent, they will encourage more thorough reviews and, ultimately, lead to safer and more effective medical tools for patients around the world.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.