Attribution in Scientific Literature: New Benchmark and Methods
This paper introduces REASONS, a large-scale benchmark and dual-metric framework designed to evaluate and improve the reliability of scientific citation attribution in large language models by balancing hallucination rates with appropriate abstention under varying evidence conditions.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the modern landscape of scientific discovery, researchers increasingly turn to artificial intelligence to summarize complex findings and connect ideas across vast libraries of literature. These intelligent systems, known as large language models, are trained on enormous amounts of text and can generate fluent, human-like responses. A critical feature of their utility in science is the ability to attach citations to their claims, pointing readers to the original papers that support a statement. However, a persistent problem undermines this trust: the models sometimes invent references that do not exist or misattribute real papers to the wrong claims. This phenomenon, often called hallucination, creates a dangerous illusion of authority. The core challenge for scientists and developers is not just to make these systems more accurate when they are confident, but to understand when they should admit they do not know. If a model cannot distinguish between a situation where it has enough evidence to answer and one where it does not, it risks confidently leading a researcher down a path of fabricated facts.
A team of researchers has addressed this uncertainty by creating a new testing ground called REASONS, a benchmark designed specifically to measure how well these models handle scientific citations. The team gathered over twelve thousand specific sentences from academic papers across twelve different fields of study, ranging from computer vision to quantum computing. Each sentence in this collection is linked to a real, verified citation, providing a clear "ground truth" against which to test the models. The researchers did not simply ask the models to find the right paper; they set up a series of experiments to see how the models behaved when the amount of information available to them changed. They tested the models in three distinct scenarios: first, with no outside information at all, forcing the models to rely solely on what they had memorized; second, with verified metadata like the paper's title and abstract provided directly to the model; and third, using advanced search tools that retrieve relevant documents from a database to help answer the question.
The results revealed a striking and consistent trade-off between being helpful and being honest. When the models were given no extra information, many of them refused to answer, a behavior the researchers call "abstention." While this refusal meant they did not make mistakes, it also meant they provided no value to the user. However, as soon as the researchers added context—either by giving the models the paper details or by letting them search a database—the models became much more willing to speak up. Unfortunately, this increased responsiveness came at a steep cost. The models stopped saying "I don't know" and started providing answers with high confidence, even when those answers were wrong. In one specific test using advanced search tools, the rate of incorrect citations jumped to nearly eighty-eight percent, while the rate of models refusing to answer dropped to zero. The models were no longer cautious; they were confidently wrong.
This behavior was not uniform across all types of science. The study found that models performed significantly worse in specialized fields like quantum computing and biomolecules compared to well-covered areas like computer vision. In these niche domains, the models were far more likely to guess incorrectly, even when provided with extra information. The researchers also tested a method where they revealed information to the model in stages, hoping that seeing more evidence would help the model correct its course. While this helped reduce some errors, it did not fix the fundamental issue: models that were prone to overconfidence simply stopped abstaining earlier and continued to assert incorrect information. The study also included a careful human review of one thousand model outputs to ensure the automated measurements were accurate. The human experts confirmed that the models were indeed making factual errors, such as citing the wrong authors or the wrong papers, rather than just using different words to say the same thing.
The findings suggest that the current approach to building these systems is missing a crucial piece of the puzzle. Simply adding more data or better search tools does not make a model more reliable; in fact, it often makes it more dangerous by suppressing its natural hesitation. The researchers argue that a truly trustworthy system must be evaluated not just on how often it gets the answer right, but on its ability to know when to stay silent. The study concludes that for artificial intelligence to be a safe partner in scientific discovery, we must design systems that value epistemic safety—the willingness to admit uncertainty—just as highly as we value their ability to generate answers. Without this balance, the risk of confident fabrication remains a significant barrier to using these powerful tools in high-stakes scientific work.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.