← Latest papers
🤖 machine learning

Robust data-driven discovery of fractional differential equations via weak formulations and Pareto-based subset selection

This paper introduces Weak-Pareto, a robust data-driven framework that combines adjoint-consistent weak formulations with Pareto-based subset selection to accurately discover fractional differential equations from noisy data by mitigating noise amplification and optimizing both discrete term structures and continuous fractional orders.

Original authors: Pongpisit Thanasutives, Yoshinobu Kawahara

Published 2026-08-14
📖 6 min read🧠 Deep dive

Original authors: Pongpisit Thanasutives, Yoshinobu Kawahara

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Detective's Dilemma: Finding Hidden Laws in a Noisy World

Imagine you are a detective trying to figure out the rules of a game just by watching the players. You see a ball rolling, a car turning, or a virus spreading, and you want to write down the exact mathematical formula that explains why it happens. This is the dream of "data-driven discovery": letting the data tell us the laws of nature instead of guessing them first.

However, real life is messy. The data we collect is never perfect; it's full of static, glitches, and random noise, like trying to hear a whisper in a rock concert. In the world of physics, some systems don't just react to what's happening right next to them; they have "long memories" or "ghostly connections" across space and time. Scientists call these "fractional" systems. Think of it like a crowd doing a wave: the person in the back doesn't just react to the person in front; they react to the whole wave pattern.

The problem is that when you try to find the math behind these ghostly connections using noisy data, the usual tools make things worse. Trying to calculate how fast something is changing (differentiation) on a messy dataset is like trying to sharpen a blurry photo by turning up the contrast—it just makes the fuzziness explode into a giant, unusable mess. This paper tackles the big question: How can we find the hidden, ghostly rules of these complex systems without letting the noise drown out the signal?


The Solution: Weak-Pareto, the Noise-Proof Detective

The researchers, Pongpisit Thanasutives and Yoshinobu Kawahara, have built a new tool called Weak-Pareto. It's a clever combination of two ideas that work together to solve the "noisy data" problem.

1. The "Smoothie" Trick (Weak Formulation)
Usually, when scientists try to find equations, they look at the data point-by-point, like checking the speed of a car at every single second. If the data is noisy, this is a disaster because the noise gets amplified, turning a tiny glitch into a massive error.

Weak-Pareto does something different. Instead of looking at individual points, it uses a "weak formulation." Imagine you have a bucket of muddy water (the noisy data) and you want to know how much dirt is in it. Instead of trying to count every single speck of dirt (which is impossible because they are too small and moving), you pour the whole bucket through a fine sieve (a smooth test function) and weigh what comes out.

In math terms, this "sieve" is a smooth, gentle curve. By multiplying the data by this smooth curve and adding everything up (integrating), the tool transfers the "roughness" from the messy data onto the smooth curve. The result? The noise gets averaged out and smoothed down, while the true signal remains clear. It's like listening to a song through a high-quality speaker system that filters out the static, rather than trying to hear it through a broken radio.

2. The "Smart Search" (Pareto-Based Subset Selection)
Once the data is smoothed out, the tool still has to guess which math terms are actually in the equation. There are millions of possible combinations of terms and numbers. A straightforward search would try every single one, but that's like trying to find a specific needle in a haystack by checking every single piece of hay one by one.

Weak-Pareto uses a "Pareto-based subset selection." Think of this as a smart shopping list. You want the best equation, but you also want it to be simple (parsimonious). If you add too many terms, the equation becomes too complicated and starts memorizing the noise instead of the rules. The tool searches for the "sweet spot"—the simplest equation that still explains the data perfectly. It balances "Loss" (how wrong the equation is) against "Complexity" (how many terms it has) to find the perfect elbow in the curve, just like finding the best price for a phone that isn't too expensive but still has all the features you need.

The Secret Sauce: Continuous Orders
Most tools have to guess the "order" of the equation from a fixed list, like choosing between 1, 1.5, or 2. But the real world doesn't always fit on a grid. The order might be 1.634. Weak-Pareto doesn't guess from a list; it searches for the exact number, treating the order as a continuous variable. This avoids the "discretization bias" where you force a round peg into a square hole.

What They Found

The team tested Weak-Pareto on several famous physics problems, including how particles spread (advection-diffusion) and how fluids move (Burgers' equation).

  • It works in the noise: When they added up to 20% noise (which is a lot of static) to the data, Weak-Pareto successfully found the correct equation in 5 out of 5 attempts for both the advection-diffusion and Burgers benchmarks.
  • The competition failed: They compared it to the old "strong-form" method (the point-by-point approach). When noise was introduced, the strong-form method failed completely, recovering the correct equation in 0 out of 5 runs. It was so overwhelmed by the noise that it couldn't find the right answer at all.
  • It handles 2D and 3D: They even showed it works in two dimensions, finding different rules for different directions (anisotropic models), which is crucial for things like how heat moves through soil or how materials stretch.
  • Speed: On the advection-diffusion benchmark, it was not only more accurate but also ran substantially faster than a neural network baseline.

The Catch and the Confidence

The authors are very careful about what they claim. They proved mathematically that for linear parts of the equation, the noise in the "weak" method vanishes as you get more data points. However, for non-linear parts (where things get messy, like u×uu \times u), the noise suppression is only "partial," meaning it helps a lot but doesn't eliminate the bias entirely.

They also found that while the tool is great at finding the structure of the equation (which terms are there), it sometimes struggles to get the exact numbers right if the noise is very high and the system is complex (like the Riesz reaction-diffusion benchmarks). In those tricky cases, it might find the right terms but get the fractional order slightly wrong, or hit the boundaries of its search range.

But the main takeaway is clear: by using the "smoothie" trick of weak formulations and a smart search for the simplest, most accurate model, Weak-Pareto can uncover the hidden laws of fractional physics even when the data is messy. It turns a noisy, chaotic signal into a clear, interpretable story about how the world works.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →