Dimensionally consistent surrogate modelling through dimensional analysis and harmonic expansions
This paper presents a data-driven method for constructing dimensionally consistent surrogate models by combining Buckingham -groups with truncated harmonic expansions, demonstrating improved conditioning, robustness, and sample efficiency across various physical systems compared to unconstrained baselines.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the physical world, the laws of nature do not care what language we use to describe them. Whether a scientist measures a distance in meters or miles, or a duration in seconds or hours, the underlying relationship between those quantities remains unchanged. This principle, known as dimensional homogeneity, acts as a strict rulebook for any equation that claims to describe reality. If a formula changes its answer simply because the units of measurement were switched, it cannot be a true law of physics. For decades, researchers have used this rule to simplify complex problems, stripping away unnecessary variables to find the core, unitless combinations that govern how systems behave. However, as scientists increasingly turn to artificial intelligence to discover these laws directly from data, a challenge has emerged. Standard machine learning models are often blind to these physical rules; they might find a pattern that fits the data perfectly but fails the moment the units change, rendering the discovery physically meaningless.
A team of researchers at the Universidad Europea de Valencia has developed a new way to teach computers to respect these physical laws from the very beginning of the learning process. Instead of letting a model guess the relationship between variables and hoping it eventually stumbles upon the correct units, they built the rules of dimensional consistency directly into the structure of the model itself. By combining classical physics techniques with modern data analysis, they created a method that forces the computer to construct its predictions using only combinations of variables that make physical sense. This approach ensures that the resulting formulas are not just mathematical tricks that work on a specific set of numbers, but robust, universal expressions that hold true regardless of how the data is measured.
The researchers tested this method on several classic problems in physics, ranging from the swinging of a simple pendulum to the radiation emitted by a perfect black body. In the case of a pendulum, the period of its swing depends on the length of the string and the strength of gravity, but not on the mass of the weight. A standard computer model, given noisy data, might struggle to figure out that mass is irrelevant or might require thousands of data points to guess the correct relationship between length and time. The new method, however, starts with the knowledge that mass should not appear in the final formula. It immediately focuses on the correct combination of length and gravity, allowing it to recover the precise physical law with far fewer data points and even when the measurements are quite noisy. The computer essentially learns the shape of the relationship much faster because it is not wasting time exploring impossible or physically nonsensical options.
The team also applied their technique to the complex radiation emitted by hot objects, a phenomenon described by Planck's law. Here, the relationship between temperature and radiation is not a simple straight line but a more intricate curve. The researchers found that by separating the problem into a part that handles the physical units and a part that learns the shape of the curve, the model could accurately reconstruct the law even with limited data. They discovered that the choice of mathematical tools used to describe the curve mattered significantly. When the data covered a range that behaved like a repeating cycle, using tools designed for waves worked best. When the data did not repeat, other mathematical tools were more effective. This nuance is crucial because it shows that while the physical rules are universal, the best way to represent the remaining details depends on the specific nature of the data.
One of the most striking findings was how much more efficient the new method was compared to standard, unconstrained machine learning. In the pendulum experiments, the new approach could accurately predict the behavior of the system with as few as five different measurements. In contrast, a standard model that was not told the physical rules had to see at least sixty measurements before it could even begin to guess the correct relationship, and even then, its predictions were less reliable. This suggests that embedding physical knowledge into the learning process is not just a way to make models more accurate, but a way to make them vastly more efficient. It allows scientists to learn from smaller, noisier datasets that would otherwise be too difficult to analyze.
The researchers also demonstrated that their method produces results that are easy for humans to understand. Unlike many modern artificial intelligence systems that act as "black boxes," hiding their logic inside layers of complex calculations, this approach yields explicit formulas. The final output is a clear equation that a physicist can read, check, and use immediately. This transparency is vital for scientific discovery, as it allows researchers to verify that the model has not just memorized the data but has actually uncovered the underlying mechanism. The study confirms that by respecting the fundamental constraints of the physical world, machine learning can become a more powerful and reliable partner in understanding the universe, turning raw data into genuine insight with greater speed and clarity.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.