Evaluation of Elicit as an adjunct search, screening, and data extraction tool for systematic reviews
This study evaluates Elicit as an AI tool for systematic reviews, finding that while it offers significant time savings and useful data extraction capabilities, its poor search overlap and low precision in screening indicate it is currently suitable only as an adjunct rather than a standalone replacement for the systematic review process.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
In the world of medical research, the systematic review stands as a cornerstone of truth. It is a rigorous process where scientists gather every available study on a specific question, sift through them with extreme care, and combine their findings to determine what is actually known about a health issue. This method is essential for doctors making treatment decisions and policymakers setting health guidelines, yet it is notoriously difficult. The process is slow, expensive, and demands immense human effort, often taking years to complete as teams of researchers read thousands of titles and abstracts to decide which papers belong in the final analysis. As the volume of published science grows every day, the sheer task of keeping up with new literature has become overwhelming, leading many to wonder if artificial intelligence could shoulder this heavy burden.
A recent study from a team at the Centre for Addiction and Mental Health in Toronto explores this very possibility by testing a specific AI tool called Elicit. The researchers wanted to know if this software could replicate the work of a human team conducting a systematic review. To find out, they took a review their own group had already completed manually—a study investigating whether early, mild psychotic symptoms in young people predict later mental health diagnoses—and asked Elicit to do the exact same job. They compared the AI's results against the human team's work at every stage, from the initial search for papers to the final extraction of data, measuring how many correct studies the AI found, how many it missed, and how much time it saved.
The results revealed a tool that is powerful but not yet ready to work alone. When it came to the initial search, Elicit struggled significantly. While the human team found over 5,000 relevant studies using traditional search methods, Elicit's search engine only uncovered 158 of those same studies. This means the AI missed the vast majority of the research that the humans had identified, suggesting that its current search capabilities are not broad enough to replace the careful, structured searches performed by human researchers. However, once the studies were in front of it, the AI showed more promise. During the screening process, where researchers decide whether a paper is relevant enough to read in full, Elicit performed moderately well. It successfully identified most of the important studies, but it also made a specific kind of error: it tended to include papers that should have been excluded. It was overly cautious, letting in conference posters and abstracts that were not full research articles, and it sometimes misjudged whether a study had the right kind of medical diagnosis data.
The most encouraging findings appeared in the data extraction phase, where the AI attempted to pull specific numbers and facts from the selected papers. Elicit proved to be a helpful assistant, successfully finding exact matches for basic details like the author's name, the year of publication, and the country of the study in nearly all cases. Even when it did not find the perfect number, it often provided a sentence from the original text that pointed the human reader to the correct information, which is a significant time-saver. For more complex details, such as the exact age of participants or specific statistical results, the AI was less precise, sometimes pulling data from the wrong group within a study or missing the specific value needed for a final calculation. Despite these imperfections, the speed difference was staggering. The human team took nearly 445 hours of work to complete the entire review process, involving multiple researchers double-checking every step. The AI completed its version of the same task in just over three hours.
Ultimately, the study concludes that Elicit is a valuable partner but not a replacement. It is not yet capable of running a systematic review on its own because it misses too many studies in the search phase and makes too many mistakes in the screening phase to be trusted without supervision. However, its ability to rapidly scan documents and point to specific sentences makes it an excellent tool for speeding up the work of human researchers. The authors suggest that the best approach is to use the AI as a second or third reviewer, letting it handle the heavy lifting of finding information while a human expert oversees the process to catch the errors and ensure the final results are accurate. In this way, the technology can reduce the time and cost of research without sacrificing the reliability that medical science requires.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.