← Latest papers
🤖 AI

Self-prompting and cross-model consensus enable reproducible data extraction from scientific literature with large language models

This paper demonstrates that combining self-prompting and cross-model consensus with large language models enables reproducible, scalable data extraction from scientific literature while maintaining expert oversight through a practical division of labor.

Original authors: Valentin Romanov, Monique Bax, Steven Niederer

Published 2026-08-20
📖 5 min read🧠 Deep dive

Original authors: Valentin Romanov, Monique Bax, Steven Niederer

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Science relies on a vast library of past discoveries, but finding the specific facts hidden inside those old papers is a slow, tedious job. Imagine a researcher needing to know exactly how a heart muscle protein behaves under certain conditions. They must read through dozens of articles, some written decades ago, looking for tiny details like temperature, chemical concentrations, and measurement methods. These details are often buried in dense text or scattered across tables, and missing even one can lead to errors in new research. For years, scientists have hoped that computer programs could do this heavy lifting, but until recently, these tools were too unreliable to trust with such delicate work. They often missed crucial context or invented facts that never existed.

A new study by researchers at Imperial College London and the University of Cambridge tests whether the latest generation of artificial intelligence can finally solve this problem. The team did not just ask a computer to read a paper; they treated the computer like a junior research assistant that needed careful training and supervision. They set up a series of experiments to see how well these programs could extract specific data from scientific articles, ranging from simple tasks where a human gave a clear set of instructions, to complex tasks where the computer had to find the articles itself. The goal was to map out exactly where these machines succeed and where they still need a human hand to guide them.

The researchers began with a straightforward challenge: could a computer extract precise numbers about how calcium binds to heart proteins from eighteen different scientific papers? They gave seven of the most advanced computer programs a detailed, expert-written set of instructions and asked them to find seven specific pieces of information for each experiment, such as the type of animal used, the temperature, and the chemical values. The results were surprisingly strong. The best programs got the numbers right more than ninety-five percent of the time. In fact, when the computers were asked to find the raw numbers alone, they were correct nearly one hundred percent of the time. However, the machines struggled when asked to understand the story behind the numbers. They often failed to identify exactly which version of the protein was being tested or the specific conditions under which the measurement was taken. This showed that while the computers were excellent at finding digits, they were still learning to understand the scientific context that makes those digits meaningful.

To see if the computers could do better without a human writing the instructions, the researchers let the programs write their own guides. They asked each computer to search for the best ways to ask questions and then to create a master instruction sheet for the task. The computers did manage to create useful guides, but the data they extracted using these self-written instructions was slightly less accurate than when they followed the human expert's original plan. The gap was small for the top programs but grew larger for others. This suggested that while computers can learn to ask good questions, a human expert still knows best how to frame the specific nuances of a scientific problem.

The team then pushed the computers further, asking them to act as autonomous researchers. In this scenario, the programs had to find the relevant scientific papers themselves, read them, and extract the data without any help from a human to point them in the right direction. This is where the systems hit a wall. Even the most advanced programs missed more than half of the correct papers they were supposed to find. Some of them invented citations for papers that did not exist, a problem known as hallucination, while others simply failed to locate the right sources. The study found that when a computer has to search for its own information, it becomes much less reliable. It is like asking a student to write a report but not giving them a library card; they might guess the right books, but they are just as likely to make up the titles.

The final part of the study looked at how to make this process work in the real world. The researchers tested a method where they used multiple computers to check each other's work. They found that if one computer made a mistake, another often caught it. By running the same task five times and taking the most common answer, the team could significantly improve the accuracy. They also discovered that having different computers work on the same problem was even better than having one computer repeat the task five times. This approach created a safety net where the computers could flag difficult cases for a human to review. The study concluded that the best way to use these tools is not to replace human scientists, but to create a partnership. In this new workflow, the computer handles the repetitive work of finding and extracting data, while the human expert steps in to verify the tricky parts and resolve any disagreements between the machines. This division of labor allows scientists to process vast amounts of literature quickly without losing the careful judgment that only a human can provide.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →