When Literature Data Mislead Artificial Intelligence in Materials Discovery
This paper demonstrates that artificial intelligence-driven materials discovery is significantly compromised by systematic errors and ambiguities in literature-derived datasets, such as text-figure mismatches and unit inconsistencies, which introduce structured label noise and necessitate improved traceable reporting and curation practices.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a world where scientists no longer spend years in the lab testing one material after another. Instead, they ask a computer to scan millions of research papers, find the best numbers, and predict which new substances will work best for batteries or solar cells. This is the promise of artificial intelligence in materials science. The idea is that if a computer can read every experiment ever published, it can learn the hidden rules of how matter behaves and design the perfect materials for our future energy needs. But this vision relies on a fragile assumption: that the numbers written in those old papers are exactly what they say they are. If a scientist writes that a material conducts electricity at a certain strength, the computer assumes that number is true, clear, and ready to use.
A team of researchers at Tohoku University and the University of Tokyo has found that this assumption is dangerously wrong. They looked closely at how scientists report data for solid electrolytes, the special materials that allow ions to move inside next-generation batteries. Their investigation reveals that the numbers extracted from scientific literature are often not just slightly off, but fundamentally misleading. The problem is not that the original scientists made mistakes, but that the way they report their findings is often ambiguous. A number might be correct in the text but wrong in the graph, or the units might be mixed up in a way that changes the value by a hundred times. When artificial intelligence systems grab these numbers to build their models, they are unknowingly learning from a distorted reality, treating vague descriptions as hard facts.
The researchers began by treating scientific papers not as finished stories, but as raw data sources. They focused on ionic conductivity, a measure of how easily electricity flows through a solid material. This value is critical for deciding if a material is good enough for a battery. In a perfect world, a researcher would measure this value, write it down, and the next person would read it and use it exactly as is. In reality, the path from a lab notebook to a computer database is full of hidden traps. The team traced thousands of reported values from their original articles into curated databases used to train artificial intelligence. They found that the journey often involved guessing what the original author meant, and those guesses were frequently incorrect.
One common issue the team found was a mismatch between the text and the figures. A scientist might write in the main paragraph that a material has a specific conductivity at room temperature, but when the researchers looked at the actual graph in the same paper, the data point for that temperature told a different story. The number in the text might be slightly higher or lower than the one shown in the plot. To a human reader, this might seem like a minor typo or a rounding error. But for a computer, which treats every number as a precise fact, this creates a conflict. The computer does not know which number to trust. When this happens across hundreds of papers, the database becomes filled with conflicting information that looks correct on the surface but is actually unreliable.
Another major source of confusion was the way graphs were labeled. Scientific papers often use complex charts to show how conductivity changes with temperature. These charts usually have axes with specific labels, but sometimes those labels are vague or missing entirely. A researcher might plot a value without clearly stating whether it represents the raw conductivity or a mathematical transformation of that value. Without this clarity, two different people reading the same graph could extract two completely different numbers. The team showed that depending on how you interpret the axis, the same data point could be read as a very high conductivity or a very low one. These differences are not random errors; they are structured mistakes that follow a pattern, making them very hard for a computer to spot as wrong.
The most dangerous errors involved the units of measurement. In science, conductivity can be reported in different units, such as per centimeter or per meter. A difference of one unit can change the value by a factor of one hundred. The researchers found cases where a paper clearly stated a value in one unit, but the graph or a later database entry treated it as if it were in another. Because the numbers still looked plausible—they were within the range of what is physically possible for these materials—the error went unnoticed. A computer reading the data would accept the wrong number as truth. In one specific case the team examined, a value was misinterpreted in a way that created a one-hundred-fold error. This single mistake, once entered into a database, was then copied and reused by other researchers and machine-learning models, spreading the error like a virus through the scientific record.
The team analyzed a large collection of papers on both gel-based and inorganic solid electrolytes. They found that text-figure mismatches were common in gel electrolytes, appearing in about ten percent of the papers they checked. For inorganic materials, the problems were even more varied, with unit inconsistencies being the most frequent issue. In a sample of sixteen inorganic cases, the researchers identified twenty-five separate instances of reporting ambiguity. This means that a single paper could contain multiple types of errors. The data showed that these were not just isolated incidents of carelessness, but a structural problem in how scientific results are communicated. Even experts in the field struggled to interpret the data correctly without making assumptions, which suggests that the problem is deep-rooted in the current system of scientific publishing.
The researchers argue that this is not just a problem for librarians or database managers; it is a crisis for the future of artificial intelligence in science. Machine learning models are only as good as the data they are fed. If the training data is full of hidden ambiguities and unit errors, the model will learn the wrong patterns. It might predict that a material is excellent when it is actually useless, or it might miss a breakthrough candidate entirely. The team found that these errors act as a form of "structured label noise." Unlike random mistakes that might average out, these errors are consistent and systematic, which means they can bias the computer's learning in a specific direction. The result is a model that appears confident but is fundamentally flawed.
To fix this, the authors propose a new standard for how scientists report their findings. They suggest that researchers must be much more explicit about the details that are currently left out. Every graph should clearly state what is being plotted and in what units. The temperature at which a value was measured should be exact, not just described as "room temperature." Authors should clarify whether a number was directly measured or calculated from a curve. These might seem like small, boring details, but the researchers insist they are the foundation of trustworthy science. Without them, the data cannot be safely reused by computers or other scientists.
The study concludes that the reliability of scientific data is an infrastructure issue, just as important as the algorithms or the computers themselves. As science moves toward a future where machines help discover new materials, the human record of experimentation must be made machine-readable in the truest sense. This means not just having numbers, but having numbers that are unambiguous, traceable, and consistent. The researchers emphasize that improving the quality of data is not a minor editorial task but a prerequisite for the next generation of scientific discovery. If the input is flawed, the output will be flawed, no matter how smart the artificial intelligence becomes. The path forward requires a shift in culture, where the clarity of reporting is valued as highly as the discovery itself, ensuring that the digital future of science is built on a solid foundation.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.