← Latest papers
🤖 machine learning

Generative Diffusion Surrogates with Analytical Variance Schedule

This paper introduces a generative diffusion surrogate for stochastic transport systems that replaces heuristic noise schedules with an analytical variance schedule derived from macroscopic theory, thereby calibrating generative time to physical transport dynamics while accurately reproducing non-Gaussian distributional features like kurtosis without requiring intermediate-time data.

Original authors: Patrick Reichherzer, Gianluca Gregori, David N. Hosking, Subir Sarkar

Published 2026-09-03
📖 7 min read🧠 Deep dive

Original authors: Patrick Reichherzer, Gianluca Gregori, David N. Hosking, Subir Sarkar

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the vast, chaotic dance of the universe, from the swirling plasma inside a fusion reactor to the jittery motion of particles in a crowded cell, there is a constant struggle between order and chaos. Scientists have long known how to describe the average behavior of these systems: if you wait long enough, a cloud of particles will spread out in a predictable way, much like a drop of ink dispersing in a glass of water. This spreading is governed by a simple rule about how far the particles travel on average over time. However, the real world is rarely that simple. The particles do not just spread; they sometimes clump, sometimes shoot ahead, and often form strange, heavy tails in their distribution that defy simple prediction. Understanding these complex shapes is crucial for everything from designing safer nuclear fusion reactors to tracking cosmic rays, yet capturing the full, messy detail of these distributions has remained a stubborn challenge for physicists.

A new approach, developed by a team of researchers at the University of Oxford and Princeton University, offers a fresh way to tackle this problem by borrowing a tool from artificial intelligence. For years, AI systems have been trained to generate realistic images by learning how to reverse a process of adding noise. Imagine taking a clear photograph and gradually blurring it until it is nothing but static; a smart computer can learn to reverse this, turning the static back into a sharp picture. The researchers realized that this same mathematical machinery could be used to model physical systems where particles spread out under random forces. But there was a catch: in image generation, the "time" it takes to blur an image is arbitrary, chosen by the programmer to make the AI work best. In physics, time is real, and the rate at which particles spread is a fundamental law of nature that cannot be faked.

The team's breakthrough was to force the AI to respect the physical clock. Instead of letting the computer decide how fast to add noise, they programmed the system to follow a specific, known law of how the particles' average spread should grow over time. They took the established physics of how particles move—whether they are bouncing through a turbulent magnetic field or drifting through a porous rock—and used that exact growth rate to control the AI's internal timer. This meant the AI no longer had to guess how fast the system was evolving; it only had to learn the shape of the distribution, the specific way the particles clustered or stretched out as they moved. By anchoring the AI to this physical reality, the researchers created a "surrogate" model that acts as a highly accurate stand-in for complex physical simulations.

To test this idea, the team focused on a scenario relevant to the future of energy: the movement of charged particles through turbulent magnetic fields, such as those found in fusion experiments or in the space around the sun. In these environments, particles do not move in a straight line; they are constantly kicked by magnetic fluctuations. The researchers used data from laboratory experiments where a beam of protons was fired through a turbulent magnetic field to create a snapshot of the turbulence. They then trained their AI model using only the starting conditions of the particles and the known physical law of how the beam should spread. Crucially, they did not feed the AI any data from the middle of the journey; the model had to learn the entire path from start to finish based solely on the beginning and the rules of spreading.

The results were striking. When the researchers asked the AI to generate the distribution of particles at a specific moment in time, the model reproduced the exact spread measured in the laboratory. It captured not just the average distance the particles traveled, but also the subtle, non-standard shapes of the distribution, including the heavy tails where rare, high-energy particles linger. The model successfully tracked how the "peakedness" of the distribution changed as the particles moved from a fast, ballistic phase into a slower, diffusive phase. This was achieved without the AI ever seeing the intermediate steps of the physical process, proving that the model had learned the underlying physics rather than just memorizing data points.

The study also revealed what happens when the physical laws are ignored. When the researchers ran the same AI model but allowed it to use a standard, arbitrary schedule for adding noise—like those used for generating pictures of cats or landscapes—the model failed to align with physical time. It could produce a blurry image that looked nice, but it did not represent the correct physical state at the correct moment. The team showed that by strictly adhering to the physical variance law, the model could act as a differentiable tool, meaning scientists could use it to work backward from a detector reading to infer the conditions at the source. This opens the door to using these models for real-time inference in complex experiments where traditional simulations are too slow or data is scarce.

One of the most significant aspects of this work is its efficiency. Traditional methods for understanding these systems often require massive amounts of data collected at every step of the journey, or they rely on solving complex equations that are computationally expensive. This new method requires only the starting distribution and the known law of spreading. It effectively separates the problem into two parts: the scale of the spread, which is handled by the known physics, and the shape of the distribution, which is learned by the AI. This allows the model to capture the complex, non-Gaussian features of the system—those heavy tails and sharp peaks that standard models miss—while remaining perfectly calibrated to the physical timeline.

The researchers validated their findings by comparing their AI surrogate against detailed computer simulations of test particles moving through a magnetic field generated by a supercomputer. The AI model matched the simulated variance and the evolution of the distribution's shape with high precision. In the specific case of the turbulent plasma, the model correctly predicted how the particles would spread over a distance of several hundred micrometers, a scale relevant to laboratory experiments. The team noted that while the model is a Gaussian surrogate in its core mechanics, it successfully inherits the non-Gaussian characteristics from the entrance data, allowing it to represent the heavy tails observed in real turbulent systems.

This approach represents a shift in how scientists might build digital twins of physical systems. Instead of trying to simulate every single collision or interaction, which is often impossible, the new method uses the known macroscopic laws to guide a generative model that fills in the microscopic details. The work suggests that for any system where the rate of spreading is known, a generative model can be calibrated to act as a precise, time-resolved emulator. This could be particularly valuable in fields like astrophysics, where direct observation is limited, or in materials science, where understanding the transport of particles through complex media is essential.

The study does not claim to solve every problem in stochastic transport. The researchers acknowledge that their model is a surrogate, a statistical approximation that works best when the underlying physical kernel is close to Gaussian or when the non-Gaussian features can be traced back to the entrance conditions. They explicitly ruled out the idea that the model could perfectly replicate every microscopic detail of a finite-speed transport process, such as the specific bounded support of a particle's path in a telegraph-like model. However, by quantifying the gap between their model and the true physical kernel, they demonstrated that the error is predictable and manageable.

Ultimately, the paper presents a method that bridges the gap between the rigid laws of physics and the flexible power of modern machine learning. By tethering the AI's internal clock to the physical reality of how particles spread, the researchers have created a tool that is both computationally efficient and physically faithful. It allows scientists to generate accurate predictions of complex transport phenomena using minimal data, turning a once intractable problem into a manageable one. The work stands as a demonstration that when artificial intelligence is grounded in the fundamental laws of nature, it can become a powerful partner in exploring the hidden dynamics of the physical world.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →