← Latest papers
🔭 astrophysics

Trustworthy scientific inference with generative models

This paper introduces FreB, a rigorous protocol that transforms potentially biased generative AI posterior distributions into statistically valid confidence regions, thereby enabling trustworthy scientific inference for complex inverse problems where traditional likelihood evaluation is infeasible.

Original authors: James Carzon, Luca Masserano, Joshua D. Ingram, Alex Shen, Antonio Carlos Herling Ribeiro Junior, Tommaso Dorigo, Michele Doro, Joshua S. Speagle, Rafael Izbicki, Ann B. Lee

Published 2026-08-21
📖 7 min read🧠 Deep dive

Original authors: James Carzon, Luca Masserano, Joshua D. Ingram, Alex Shen, Antonio Carlos Herling Ribeiro Junior, Tommaso Dorigo, Michele Doro, Joshua S. Speagle, Rafael Izbicki, Ann B. Lee

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Scientists have long faced a fundamental puzzle: how to learn the hidden rules of the universe when they can only see the results. Imagine looking at a shadow on a wall and trying to figure out the exact shape of the object casting it. In fields ranging from astronomy to particle physics, researchers observe data—light from distant stars, collisions of subatomic particles, or patterns in the atmosphere—and must work backward to deduce the properties that created them. This is known as an inverse problem. For decades, scientists relied on complex mathematical formulas to make these deductions, but as our instruments have become more powerful, generating vast amounts of high-resolution data, those old formulas have often become too slow or too difficult to use. In response, many have turned to generative artificial intelligence. These are computer systems trained on examples that can learn to predict the hidden causes of observed data with incredible speed. They have become essential tools for analyzing everything from the James Webb Space Telescope's images of the early universe to the chaotic spray of particles in underground detectors.

However, a critical flaw has emerged in this new approach. While these artificial intelligence models are excellent at guessing the most likely answer, they often fail to tell scientists how sure they should be. When a model is trained on a specific set of examples, it can become overconfident, offering narrow, precise-sounding answers that are actually wrong, especially when the real-world data it encounters looks slightly different from its training. This is dangerous for discovery. If a scientist believes a measurement is precise when it is not, they might draw false conclusions about the existence of new particles or the history of a galaxy. The core challenge is not just making a prediction, but ensuring that the range of possible answers provided by the computer actually contains the true value with a known, reliable probability, no matter what the true value happens to be.

A team of researchers has now developed a solution to this problem, introducing a new protocol called Frequentist-Bayes, or FreB. This method acts as a rigorous quality-control check for the outputs of generative AI. Instead of simply accepting the AI's initial guess, FreB reshapes the computer's probability map into a statistically trustworthy confidence region. The process begins by training the AI on a large set of labeled examples, where both the hidden properties and the resulting data are known. The AI learns to create a "posterior" distribution, which is essentially a map showing where the true answer is likely to be. The researchers then use a separate set of calibration data to teach the system how to adjust this map. They transform the AI's raw probabilities into what are known as p-value functions. These functions act like a calibrated ruler, ensuring that when the system draws a boundary around a set of possible answers, that boundary will contain the true value exactly as often as the scientists claim it will.

The power of this approach lies in its ability to handle real-world imperfections. In many scientific studies, the data used to train the AI does not perfectly match the data being studied. This mismatch, known as dataset shift, can happen because of how telescopes are built, which stars they can see, or because the theoretical models used for training are slightly inaccurate. Traditional AI methods often produce biased results in these situations, leaning too heavily on the training data. FreB, however, is designed to correct for these differences. By using the calibration data to learn the necessary adjustments, the method ensures that the final confidence regions remain valid even when the training data and the target data come from different distributions. The researchers demonstrated that this works even when the AI model itself is slightly misspecified, meaning the underlying physics model used to generate the training examples was not a perfect representation of reality.

To prove their method works, the team tested it on three distinct challenges in the physical sciences. In the first case, they tackled the reconstruction of gamma rays, which are high-energy particles from space that cannot be seen directly but are inferred from the showers of secondary particles they create in Earth's atmosphere. The AI was trained on data from the Crab Nebula, a well-known source, but was then asked to identify signals from a dark matter source, which produces very different, rarer patterns. Standard AI methods failed here, producing estimates that were biased toward the energy levels of the training data and missing the true, higher-energy events. The FreB method, however, successfully reshaped the AI's output to provide valid confidence intervals that correctly captured the true energy of the rare events, allowing scientists to distinguish between different astrophysical sources with reliability.

The second test involved understanding the history of our own galaxy, the Milky Way. Astronomers use different theoretical models to describe how stars are distributed and how they move. These models can sometimes lead to conflicting conclusions about a star's age or chemical makeup. When the researchers applied standard AI inference, the different models produced results that disagreed sharply with each other and with the true values, offering no clear way to resolve the tension. By applying FreB, the team was able to adjust the outputs of both models. The result was a set of confidence regions that, despite coming from different starting assumptions, both correctly contained the true properties of the star. This showed that FreB could reconcile competing theories, ensuring that the uncertainty estimates were honest and consistent regardless of the underlying model used.

The final challenge addressed the issue of selection bias, a common problem in large astronomical surveys. Telescopes often cannot see every star equally; they are better at detecting bright, large stars than faint, small ones. This means the data used to train AI models is often skewed, missing the very objects scientists most want to study. In this experiment, the researchers simulated a scenario where the AI was trained on data dominated by large, bright stars, but was then asked to estimate the properties of smaller, fainter stars like our Sun. The unadjusted AI produced estimates that were systematically wrong, failing to cover the true values for these smaller stars. FreB corrected this bias by using a small set of follow-up data to recalibrate the system. The adjusted method produced confidence regions that were both tight and accurate, successfully covering the true parameters of the faint stars even though the training data had been heavily biased against them.

The significance of this work extends beyond these specific examples. It provides a mathematical guarantee that the confidence intervals generated by these powerful AI tools are trustworthy. In the past, scientists had to hope that their models were accurate enough to provide valid uncertainty estimates. Now, they have a protocol that can take a fast, flexible AI model and reshape its output to meet the strict standards of scientific proof. This means that in fields where direct calculation is impossible or too expensive, researchers can use generative AI to explore complex systems with the same level of statistical rigor they would expect from traditional methods. The method is efficient, requiring no retraining when applied to new data, and it works even when the training data is imperfect. By bridging the gap between the speed of modern artificial intelligence and the reliability of classical statistics, FreB offers a path toward more trustworthy scientific discovery, ensuring that when scientists look at the shadows on the wall, they can be confident in the shape of the object casting them.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →