BootLoops: an LLM-driven toolkit for exact quantitative science
BootLoops is an LLM-driven toolkit that unifies advanced exact computation methods from physics, mathematics, and computer science to enable agentic models to perform rigorous quantitative analyses—such as calculating Feynman integrals, Bayesian evidence, and configuration enumerations—across diverse scientific fields with guaranteed precision.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
For most of the history of scientific computing, a human being would write a program, knowing exactly what it does, and then a computer would carry it out. Later, the paradigm shifted: a human would specify the architecture, a computer would fit a neural network to data, and the human would interpret the results. Now, a new shift is underway. A large language model, acting as a semi-autonomous agent, can design and write the program, run it, and then review and interpret the answer itself. The potential of such systems seems unbounded, yet a key question remains: how can humans be maximally productive using these agents? The answer lies in connecting distant fields. Methods developed for particle physics might solve problems in ecology, or tools from numerical analysis could clarify regulatory thresholds in public health, but few people know both sides well enough to make the connection. A large language model, having read everything, knows both.
This paper introduces BootLoops, a toolkit designed to harness that breadth of knowledge. It is a collection of specialized computer programs that a language model can operate and extend to perform exact, quantitative science. The toolkit brings together methods from collider physics, experimental mathematics, and computer algebra to solve problems that are usually estimated with approximations or computed with floating-point numbers that hide rounding errors. Instead of guessing or estimating, BootLoops allows the model to compute answers as exact rational numbers, prove that a solution exists or does not exist, and carry calculations with guaranteed error bounds. The goal is not to replace human scientists, but to give them a way to use tools from one field to solve problems in another, turning connections that are hard for humans to make into routine operations.
The core of the work involves a large language model running in a loop, issuing commands to execute programs and reading their output. When faced with a difficult calculation, the model plans the steps, selects the right tools from the toolkit, and runs them. If a tool is missing, the model writes a new routine, tests it, and adds it to the collection. The programs themselves compute every number and symbol; the model never guesses a digit. The toolkit persists across different problems, growing as new methods are added. It includes programs that reduce complex integrals to a finite basis, solve differential equations, and fit exact constants to high-precision numbers. Some of these programs are public codes adopted from physics and mathematics, while others were written from scratch specifically for this project. In total, roughly one hundred new programs were written just to handle the computations described in the paper.
The first major application of BootLoops is the calculation of Feynman integrals, which describe how particles scatter and interact. For decades, physicists have computed these by direct integration or by estimating them point by point. BootLoops takes a different approach, using a method called a "bootstrap." The model first identifies the geometric shape defined by the problem, which could be a simple sphere, an elliptic curve, or a complex shape known as a Calabi-Yau manifold. This geometry dictates the type of functions that can describe the answer. The model then builds a general formula with unknown coefficients and uses high-precision numerical samples to fit those coefficients. The result is a closed-form answer, a precise mathematical expression, rather than a list of numbers. The author computed thirty such integrals, ranging from simple one-loop diagrams to complex four-loop structures. Fifteen of these were known in the literature, and the other fifteen were new, with no prior computation found. Each result was checked against an independent numerical evaluation to thirty or more digits of precision, ensuring the answer was correct without relying on the model's intuition.
The second application shows how these physics tools solve problems in biology and statistics. In fields like evolutionary biology, scientists compare different family trees of species to see which one best explains the DNA data. This requires calculating a "Bayesian evidence," a number that represents the probability of the data under a specific tree. Traditionally, this is estimated using Monte Carlo methods, which are slow and come with their own estimated errors. BootLoops recognizes that the math for these biological evidence integrals is identical to the math for Feynman integrals. By applying the same reduction and fitting techniques, the toolkit computes these evidences as exact rational numbers. For simple cases, the answer is a precise fraction. For more complex cases, it provides a guaranteed range that contains the true value. This allows scientists to compare trees with absolute certainty, something that was previously impossible with floating-point estimates. The same approach is applied to mixture models in statistics, population genetics, and single-cell biology, where it provides exact answers for problems that usually rely on noisy simulations.
The third capability is the ability to perform an exhaustive search with a proof of completeness. In many scientific questions, the answer depends on checking every possible case, or proving that no case of a certain kind exists. BootLoops can enumerate every configuration in a finite space, such as the possible arrangements of flux in string theory or the optimal ways to group data for quality ratings. It does this with a rigorous proof that nothing was missed. If the search finds a solution, it returns the complete list. If it finds no solution, it provides a theorem proving that no such object exists. For example, the toolkit was used to analyze the Watson integral, a famous problem in mathematics regarding random walks. By exhausting the class of possible closed-form solutions, the model proved that no such closed form exists for the general case, a result that had been an open question for decades. This transforms a search that might have been heuristic or incomplete into a mathematical certainty.
Finally, the toolkit addresses the problem of rounding errors in standard calculations. When computers use floating-point arithmetic, tiny errors accumulate and can lead to wrong answers, especially in long chains of calculation. BootLoops uses "ball arithmetic," where every number is treated as a midpoint with a guaranteed radius of error. This radius grows or shrinks as calculations proceed, ensuring that the true value is always contained within the bounds. If the error becomes too large to be useful, the system refuses to give an answer rather than giving a wrong one. This method was used to audit published numbers, such as regulatory thresholds for air quality and medical ratings. In one case, the toolkit replayed the calculation for a US ozone standard using exact arithmetic and found that the published result depended on how the numbers were truncated, a detail that floating-point arithmetic had obscured. In another, it verified the exact values for Medicare star ratings, proving that the published cut points were indeed optimal.
The paper concludes that the true power of agentic artificial intelligence lies in its ability to combine knowledge from different fields. BootLoops is a first attempt to pool these methods into a single, extensible toolkit that a language model can operate. It does not push the boundaries of human knowledge in a single direction but rather fills in the gaps between existing fields. The toolkit is open source and designed to grow, with the expectation that scientists will add new methods as they are discovered. The author emphasizes that while the model writes the code and runs the tools, the results are not based on the model's judgment but on the deterministic execution of the programs. The work demonstrates that by bringing together the reduction algorithms of collider physics, the evidence integrals of statistics, and the rigorous bounds of numerical analysis, we can solve long-standing problems with a level of precision and certainty that was previously out of reach. The toolkit and its documentation are available for anyone to use and extend, marking a step toward a future where the convex hull of human scientific knowledge is fully explored.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.