← Latest papers
💬 NLP

Provenance Before Prose: Claim-Locked Reporting

This paper introduces "claim-locked reporting," a provenance-before-prose protocol that fixes statistical evidence and claim parameters before language generation, significantly improving numerical reproducibility and reducing computational costs compared to existing hybrid template methods.

Original authors: Xiao Fan, Jingyuan Li, Hongbin Guo, Yubo Han, Yi Zhang

Published 2026-08-27
📖 6 min read🧠 Deep dive

Original authors: Xiao Fan, Jingyuan Li, Hongbin Guo, Yubo Han, Yi Zhang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the world of modern science, researchers often rely on powerful computer programs to turn raw data into written reports. Imagine a scientist who has spent months running complex experiments to see how the human brain changes with weight, or how a new drug affects a disease. Once the math is done, the results are fixed numbers: a specific count of connections, a precise measurement of change, and a clear direction of effect. Today, scientists increasingly ask artificial intelligence to take these fixed numbers and write the narrative that explains them to the public or to doctors. The hope is that these computer programs can write fluently and quickly. However, there is a hidden danger in this process. Even when the starting numbers are correct, the computer can sometimes change the story it tells. It might accidentally swap a number, flip the direction of an effect, or describe a weak finding as a strong one. This happens because the computer is not just writing; it is also making choices about which facts to highlight and how strongly to state them, and those choices can vary every time the program runs.

Researchers at Xidian University have developed a new way to stop this from happening, ensuring that the story the computer writes stays perfectly locked to the facts. They call this method "claim-locked reporting." Instead of letting the computer decide which numbers to use or how to phrase them, the researchers force the computer to first agree on a strict list of claims. For every single point the report will make, the system first checks the original data to confirm the exact number, the direction of the effect, and how strong the evidence really is. Only after this list is finalized and locked in does the computer begin to write the connecting sentences. Think of it like a construction crew that must lay every brick exactly where the blueprint says before they ever start mixing the mortar to hold them together. In their study, the researchers tested this approach on two very different types of scientific writing: reports on brain imaging and summaries of clinical drug trials. They found that when they used this new method, the computer produced the exact same numbers and facts every single time it ran, whereas older methods changed the details frequently.

The team tested their idea by asking the computer to write reports based on real data from brain scans of people with obesity and from public records of medical drug trials. In the brain scan tests, they compared their new method against several older ways of using artificial intelligence. The older methods, which included simple instructions or templates that let the computer pick its own numbers, produced reports that were only about 61 percent consistent with each other. This means that if you ran the same experiment twice, the computer would often give you different numbers or different conclusions. In contrast, when the researchers used their claim-locked method, the reports were 98.5 percent consistent. The computer no longer chose which facts to include; it simply reported the facts that had already been approved. This consistency was not just about repeating numbers; it also meant the computer stopped making dangerous mistakes, such as describing a weak link as a strong one or reversing the direction of a medical effect.

The study also looked at how well the computer handled tricky situations where the meaning of a number depends on other factors. In one specific test involving brain scans, the data showed a large number of changes when looking at groups of people, but when the researchers adjusted for body weight, that number dropped almost to zero. This meant the initial large number was misleading if reported without the adjustment. Older methods often failed here, reporting the large number as if it were a solid, independent fact. The new method, however, caught this. Because the system locked the claim to the full context of the data before writing, it correctly reported that the large number disappeared once weight was taken into account. It described the result with the right level of caution, avoiding the mistake of claiming a direct cause where none existed.

To make sure these results were not just a result of the computer code, the researchers also asked human experts to read the reports. These experts were blind to which method had written each report. They looked for errors in direction and for sentences that were too strong or unsupported. The human reviewers confirmed that the new method produced far fewer errors. In the drug trial tests, the older methods sometimes flipped the direction of a result, saying a drug helped when it actually hurt, or vice versa. The new method never made this mistake; it preserved the correct direction every time. The human auditors also noted that the new method was much better at using cautious language when the evidence was weak, whereas the older methods tended to sound overly confident.

Beyond accuracy, the new method also turned out to be more efficient. Because the computer did not have to spend time deciding which numbers to use or rewriting sections to fix errors, it used fewer computing resources and finished the work faster. In the tests, the new method required less time and less computer power to generate a report than the older methods that tried to check their work after writing it. This suggests that building the facts in first is not only more reliable but also cheaper and faster.

The researchers are careful to note that their method does not fix bad data or bad science. If the original numbers are wrong or the experiment was flawed, the report will still reflect those errors. The system simply ensures that the computer does not add its own mistakes on top of the existing ones. It acts as a strict gatekeeper for the facts, ensuring that the story told matches the evidence exactly. By moving the control from the writing phase to the planning phase, the team has shown that it is possible to get artificial intelligence to write scientific reports that are both fluent and rigorously faithful to the data. This approach offers a practical path forward for using computers to communicate science without losing the precision that scientific discovery requires.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →