← Latest papers
💻 computer science

ARISMA: Guidelines for AI- and LLM-Assisted Systematic Reviews, Scoping Reviews, and Mapping Studies

This paper introduces ARISMA, a comprehensive guideline framework that establishes principles, a lifecycle taxonomy, and reporting standards for integrating AI and large language models as auditable, human-supervised assistants in systematic reviews, scoping reviews, and mapping studies to ensure methodological rigor and accountability.

Original authors: Mahyar Tourchi Moghaddam, Mina Alipour

Published 2026-08-27
📖 6 min read🧠 Deep dive

Original authors: Mahyar Tourchi Moghaddam, Mina Alipour

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the world of scientific research, there is a specific kind of work known as a systematic review. Imagine a team of experts trying to answer a complex question, such as whether a new drug works better than an old one, or what the current state of knowledge is about climate change. To do this, they cannot simply read a few interesting articles; they must find every single relevant study ever published on the topic, read them all, and combine their findings into a single, trustworthy conclusion. This process is the foundation of modern medicine and policy, but it is incredibly difficult. As the number of scientific papers grows every year, the task of finding and reading them all has become slower, more expensive, and harder to keep up to date.

Recently, powerful computer programs known as artificial intelligence have entered the scene. These tools can read text, understand language, and summarize information at speeds no human can match. Researchers began using them to help find papers, decide which ones to keep, and pull out important numbers. However, this new speed came with a problem. No one had a clear set of rules for how to use these tools safely. If a computer program missed a crucial study or made a mistake in reading a number, the entire review could be wrong, and the doctors or policymakers relying on it could make bad decisions. The question became not just whether these tools could work, but how to use them without losing control of the science.

A team of researchers from the University of Southern Denmark has now proposed a new set of guidelines called ARISMA to solve this problem. Their work does not treat artificial intelligence as a replacement for human scientists. Instead, they frame it as a highly skilled assistant that must be watched, tested, and logged at every step. The core idea is simple but strict: every important decision in a scientific review must remain understandable, checkable, and ultimately the responsibility of a human being. The researchers argue that while machines can do the heavy lifting, they cannot be allowed to silently decide what counts as evidence.

The paper outlines a complete lifecycle for how a review should be conducted when using these tools. It begins with the planning stage, where the human team must clearly define what they are looking for before any computer is involved. The guidelines suggest that artificial intelligence can be very helpful in brainstorming search terms or suggesting synonyms for a topic, but the final search strategy must be checked by a human expert. The researchers emphasize that if a computer misses a known, important study during a test run, the search strategy is considered a failure and must be fixed before it is used on the real data. This ensures that the search is not just fast, but also complete.

Once the search is done, the team faces the massive task of screening thousands of titles and abstracts to decide which papers to read in full. This is where artificial intelligence has shown the most promise, but also where the guidelines are most cautious. The paper explains that a computer can be used to rank papers or suggest which ones to exclude, but it cannot make the final cut on its own. The researchers propose a system of "trust levels." In a low-trust scenario, the computer only suggests an order for humans to read. In a moderate-trust scenario, it might recommend exclusions, but a human must check every single one of those recommendations. In high-stakes reviews, such as those that will guide medical treatments, the computer acts only as a second pair of eyes, and the human team makes the final call. Crucially, before a computer is allowed to do any of this, it must be tested on a small set of papers that humans have already graded. If the computer makes too many mistakes on this test set, it is not allowed to proceed.

The guidelines also address the tricky task of data extraction, where researchers pull specific numbers like patient counts or treatment results from the text of a paper. Here, the researchers are even more conservative. They note that while computers are good at finding general descriptions, they often struggle with exact numbers hidden in tables or complex sentences. The paper suggests that a computer can fill in a draft form with descriptive details, but every single number that will be used to calculate a final result must be checked line-by-line by a human. If a human does not verify the source of a number, that number cannot be used in the final review. This rule protects against the computer "hallucinating" or inventing facts that sound plausible but are not in the original text.

Throughout the process, the guidelines demand a paper trail. Every time a computer tool is used, the researchers must record what tool was used, what version of the software it was, what instructions were given to it, and what the output was. If the computer makes a mistake, or if the human team decides to change how the computer is working, that change must be dated and explained. This creates a history of the review that allows anyone reading the final report to see exactly where the computer helped and where the humans took over. The paper also touches on the legal and ethical side, reminding teams that they must be careful about sending private or unpublished data to public computer services, as that data could be stored or used in ways the team did not intend.

The researchers developed these guidelines by talking to twenty-one experts in the field, including scientists and information specialists. They tested their ideas and refined them based on feedback, focusing on making the rules practical and clear. The result is a framework that allows teams to use the speed of artificial intelligence without sacrificing the accuracy and trustworthiness that scientific reviews require. The paper concludes that the goal is not to make reviews fully automatic, but to make them more auditable. By treating artificial intelligence as a tool that is inspected and bounded rather than an autonomous decision-maker, the guidelines ensure that the final conclusions of a review remain grounded in verified evidence, with a human always in the loop to take responsibility for the truth.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →