← Latest papers
⚛️ high-energy experiments

POPxf: An Exchange Format for Polynomial Observable Predictions

The paper introduces POPxf, a structured, machine-readable data format designed to encode semi-analytical theoretical predictions as polynomials in model parameters—particularly for Effective Field Theory applications—to enhance reproducibility, facilitate global fits, and streamline the exchange of uncertainty-aware results within the high-energy physics community.

Original authors: Ilaria Brivio (ed.), Ken Mimasu (ed.), Peter Stangl (ed.), Anke Biekötter, Ana R. Cueto Gómez, Charlotte Knight, Luca Mantani, Eleonora Rossi, Alejo N. Rossia, Aleks Smolkovič

Published 2026-10-01
📖 7 min read🧠 Deep dive

Original authors: Ilaria Brivio (ed.), Ken Mimasu (ed.), Peter Stangl (ed.), Anke Biekötter, Ana R. Cueto Gómez, Charlotte Knight, Luca Mantani, Eleonora Rossi, Alejo N. Rossia, Aleks Smolkovič

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the vast, invisible landscape of particle physics, scientists act as cosmic detectives, trying to understand the fundamental rules that govern the universe. They do this by smashing particles together at incredible speeds, creating a shower of debris that reveals the existence of forces and particles too small to see directly. To make sense of these collisions, physicists rely on mathematical models that predict what should happen if their theories are correct. For decades, a powerful tool called the Effective Field Theory has allowed them to describe these interactions by focusing on the most important variables while ignoring the messy details of the unknown. However, a significant bottleneck has emerged: while these theories are powerful, the actual predictions they generate are often locked away in complex computer code or hidden within dense mathematical papers. This makes it difficult for different research teams to share their findings, check each other's work, or combine their results to get a clearer picture of reality.

A new proposal from a collaboration of theorists at CERN and universities across Europe aims to break down these barriers. They have introduced a standardized, machine-readable format called POPxf, designed to act as a universal translator for theoretical predictions. Instead of hiding their calculations in proprietary software, researchers can now package their predictions in a simple, structured file that anyone can read and use. This format captures the essence of how physical quantities change as the underlying parameters of the universe shift, allowing scientists to exchange complex data with the same ease as sharing a spreadsheet. By making these predictions transparent and accessible, the group hopes to streamline the process of testing the Standard Model of particle physics and searching for new physics beyond it.

The core of this work is the recognition that many predictions in particle physics can be described as polynomials. In everyday terms, a polynomial is a mathematical expression built from a set of numbers and variables combined through addition and multiplication. In the context of high-energy physics, these variables represent the strength of interactions between particles, known as Wilson coefficients. When a particle collision occurs, the probability of a specific outcome—such as a particle decaying into a specific set of other particles—can be calculated by plugging these coefficients into a polynomial equation. This structure is incredibly common, appearing in analyses of the Higgs boson and in searches for new physics, yet it has historically been difficult to share the actual coefficients that define these equations.

The researchers behind this note have developed a specific data format to solve this problem. They chose JSON, a lightweight text-based format that is already widely used in web development and is easily readable by both humans and computers. The format is designed to be flexible enough to handle simple cases where a prediction is a single polynomial, as well as more complex scenarios where an observable is a function of several polynomials, such as a ratio of two different decay rates. Crucially, the format does not just store the final numbers; it records the entire mathematical structure, including the specific variables used, the scale at which the calculation was performed, and the assumptions made about the underlying physics. This level of detail ensures that anyone using the data can reproduce the original result exactly, eliminating the guesswork that often plagues scientific collaboration.

One of the most significant features of this new format is its ability to handle uncertainties and correlations. In any scientific measurement, there is always a margin of error, and these errors are often linked. For instance, if two different measurements depend on the same theoretical input, an error in that input will affect both results in a similar way. The POPxf format allows scientists to explicitly record these relationships, distinguishing between uncertainties that are constant and those that change depending on the values of the model parameters. This is a vital improvement over previous methods, which often treated uncertainties as fixed numbers, potentially leading to inaccurate conclusions when the parameters of the model were varied. By capturing how errors shift and interact, the format provides a more honest and complete picture of the reliability of the predictions.

The proposal also addresses the issue of metadata, or the "data about the data." A prediction is useless without knowing the context in which it was created: what computer program was used, what version of the software, what physical constants were assumed, and what specific mathematical approximations were made. The new format includes dedicated fields to record all of this information, creating a digital footprint that allows future researchers to trace the origin of a result. This is particularly important for the global effort to combine data from different experiments, such as those at the Large Hadron Collider, into a single, coherent analysis. Without a standard way to document these details, combining results from different teams is like trying to assemble a puzzle where the pieces are from different boxes and the picture on the box is missing.

To demonstrate how this works in practice, the authors provide several examples, including predictions for the decay rates of the W boson and the branching ratios of B mesons. In these examples, the format successfully encodes the central values of the predictions, the associated uncertainties, and the complex correlations between them. The data is organized in a way that allows software tools to automatically read the file, extract the relevant numbers, and plug them into larger statistical analyses. This automation is key to the format's success, as it removes the need for researchers to manually transcribe data, a process that is prone to human error and time-consuming.

The development of this format represents a shift in how theoretical physics is communicated. It moves away from the traditional model of publishing a static paper with a few key results and toward a dynamic ecosystem where data is shared, validated, and reused. The authors have made the specifications for this format publicly available, along with a collection of example files and tools to help others adopt it. They envision a future where theoretical predictions are as easy to share and verify as experimental data, fostering a more collaborative and efficient scientific community. By providing a common language for the mathematical descriptions of the universe, this work aims to accelerate the discovery of new physics, ensuring that the collective effort of the global community is not held back by incompatible data formats.

The paper does not claim to have solved every problem in data exchange, nor does it suggest that all theoretical predictions can be perfectly captured in a single file. It acknowledges that some complex scenarios may require additional tools or formats, and it leaves room for the standard to evolve as the field advances. However, it firmly establishes that a standardized, machine-readable approach is necessary to overcome the current fragmentation in the field. The authors argue that without such a standard, the community risks duplicating effort and missing subtle connections between different measurements. By adopting this new format, they believe the field can move forward with greater confidence, knowing that the theoretical foundations of their experiments are built on a shared, transparent, and reproducible base.

Ultimately, this work is about clarity and connection. It is an attempt to bring the abstract world of theoretical predictions into the concrete realm of shared data, making the invisible machinery of the universe more accessible to those who study it. By turning complex mathematical relationships into structured, readable files, the researchers have provided a tool that empowers scientists to focus on the physics rather than the logistics of data handling. In a field where the stakes are high and the questions are profound, this simple act of standardization could prove to be a powerful catalyst for discovery, ensuring that the insights gained from the world's most powerful particle accelerators are fully understood and utilized by the entire scientific community.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →