← Latest papers
💻 computer science

CompProv Produces Machine Readable Graphs Encoding Microscopic Algebraic Provenance for Reproducible Computation

This paper introduces CompProv, a Java-based framework that ensures reproducible and auditable computational results by capturing microscopic algebraic provenance at the atomic operation level through a serializable Calculation Provenance Graph, enabling deterministic replay and sensitivity analysis without exposing proprietary source code.

Original authors: Minas Abramyan, Mohammed Alaa Ala'anzy, Nasir Saeed

Published 2026-10-02✓ Author reviewed ⓘ
📖 5 min read🧠 Deep dive

Original authors: Minas Abramyan, Mohammed Alaa Ala'anzy, Nasir Saeed

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the world of scientific research and financial modeling, trust has traditionally been a matter of faith. A researcher publishes a result, or a bank reports a portfolio value, and the audience must assume the numbers were calculated correctly. For decades, the tools used to verify these numbers have operated like a shipping manifest: they can tell you which box of data arrived at the factory and which box left, but they cannot see what happened inside the factory floor. If a calculation involves a long chain of mathematical steps, a tiny rounding error or an undocumented change in a number can slip through the cracks, silently altering the final answer. Once the calculation is finished and the computer moves on, those intermediate steps vanish, leaving no trace of how the result was actually derived. This gap has become a major obstacle for anyone trying to verify complex work, from climate scientists checking weather models to auditors verifying billions of dollars in assets.

A team of researchers has developed a new system called CompProv that changes how we look at these calculations. Instead of just watching the start and end points, this system records every single mathematical move a computer makes, step by step. It works by wrapping every number in a special digital container that automatically writes down its own history. As the computer adds, subtracts, or multiplies these numbers, the system captures the exact operation, the time it happened, and where the numbers came from. The result is a complete, self-contained map of the entire calculation that can be saved, shared, and replayed years later, even if the original software or computer environment no longer exists.

The researchers tested this idea in three very different fields to see if it could hold up in the real world. First, they applied it to a decentralized finance scenario, simulating the calculation of a portfolio's total value. In this test, the system tracked the value of various digital assets as they were converted and combined. The team found that they could take the final record, feed it into a fresh computer environment, and get the exact same result every time. More importantly, they could swap out specific input numbers, such as changing the price of a single asset, to see how the final total would shift, all without ever needing to see the original, private code that performed the calculation. This proved that the system could provide a clear audit trail for sensitive financial data without exposing the proprietary logic behind it.

Next, the team turned to the strict world of physical measurement, specifically the calibration of a gauge block used to define precise lengths. In this field, the chain of evidence must be unbroken to meet international standards. The researchers reconstructed a published calibration procedure, but they discovered a significant problem: the original paper had left out seven of the thirteen numbers needed to perform the calculation, such as air temperature and pressure. Because the original authors had not recorded these intermediate values, the calculation could not be fully reproduced from the published text alone. The new system, however, made these missing pieces visible. By forcing every number to be tagged with its source before it could be used, the system highlighted exactly which values were missing or had to be assumed. The resulting record showed not just the final length, but the entire chain of assumptions and measurements that led to it, making the limitations of the original study transparent and auditable.

The final test involved a hydrological model used to predict river flow. The researchers took a complex simulation that had been run previously and used the system to re-calculate its performance metrics. In the original study, the final scores were reported, but the detailed steps showing how those scores were derived were not included. The new system captured the entire process, from the raw water flow data to the final performance score. The researchers were able to replay the calculation and verify the results exactly. They also used the system to run a sensitivity analysis, substituting different input data to see which simulation performed best. This allowed them to identify the most accurate model and prove that the result was not a black box, but a fully traceable chain of events tied directly to the original data.

To ensure the system could handle the demands of real-world computing, the researchers also measured how much extra work it added to the process. They found that for simple, fast calculations, the system did slow things down, as it had to write down every step. However, they developed a technique to compress the record for repetitive loops, which significantly reduced the extra time and memory required. Even with this overhead, the system proved capable of running complex tasks in parallel, scaling up to use multiple computer processors just as efficiently as standard software. The study showed that while the system requires more resources than a standard calculation, it provides a level of verification that was previously impossible.

The work demonstrates that it is possible to build a system where the history of a number is as important as the number itself. By making the internal steps of a calculation visible and permanent, the researchers have created a tool that allows scientists and auditors to verify results without needing to trust the original software or the original environment. The system does not fix errors that happen during a calculation, but it ensures that if an error occurs, or if a number was changed, there is a complete, unalterable record of exactly what happened. This shifts the standard of proof from trusting a reported output to verifying a documented process, offering a new way to ensure integrity in an increasingly complex digital world.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →