← Latest papers
🔬 physics

Missing the Butterfly and Predicting the Past: Features or Bugs of Accurate AI Weather Models?

This paper reveals that the high forecast accuracy of AI weather models, along with their ability to skillfully backcast and their failure to exhibit the butterfly effect, stems from the inevitable coarse-graining of training data, which inadvertently filters out fast, small-scale error growth while preserving large-scale patterns.

Original authors: Pedram Hassanzadeh, Weidong Li, Y. Qiang Sun, Jiangdi Wang, Alexander Wikner, Justin Finkel, Jonathan Q. Weare

Published 2026-08-27
📖 6 min read🧠 Deep dive

Original authors: Pedram Hassanzadeh, Weidong Li, Y. Qiang Sun, Jiangdi Wang, Alexander Wikner, Justin Finkel, Jonathan Q. Weare

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Weather prediction has long been a battle between two opposing forces: the chaotic nature of the atmosphere and the limits of our ability to measure it. The atmosphere is a fluid system that changes constantly, where a tiny shift in conditions today can lead to a completely different storm weeks later. This sensitivity is often called the butterfly effect, a concept suggesting that the flap of a butterfly's wings could eventually alter the path of a hurricane. Because of this, traditional weather models, which rely on complex physics equations to simulate the air, struggle to predict the future with perfect accuracy. If the starting data is even slightly off, the errors grow rapidly, making long-term forecasts unreliable.

In recent years, a new type of weather model has emerged, powered by artificial intelligence. These systems do not solve physics equations; instead, they learn patterns from decades of historical weather data. They have proven surprisingly accurate, often outperforming the best traditional models while using a fraction of the computing power. However, this success has raised a puzzling question for scientists: how can these models be so good at predicting the future if they seem to ignore the very chaos that makes weather so hard to predict? If they truly understood the atmosphere's physics, they should struggle with the butterfly effect just like traditional models do. Yet, they do not.

A team of researchers set out to solve this mystery by testing these artificial intelligence models in a series of controlled experiments. They wanted to see if the models were truly learning the laws of physics or if they were simply finding a shortcut that worked for the future but failed the test of time. To do this, they trained the models not just to predict the weather forward, but also to predict it backward. In the real world, predicting the past from the present is nearly impossible for a chaotic system because the process of dissipation—where energy is lost as heat—cannot be easily reversed. It is like trying to un-mix a cup of coffee and cream; the laws of physics make it impossible to separate them once they are blended.

The researchers found that the artificial intelligence models could indeed predict the past with surprising skill. They could look at a weather map from today and accurately reconstruct what the atmosphere looked like days ago. This result was startling because it seemed to violate the fundamental rules of thermodynamics, which dictate that time moves in one direction and that reversing it should lead to chaos. However, the models were not breaking the laws of physics; they were bypassing them. The study revealed that the models were trained on data that had been smoothed out, or "coarse-grained." This process removed the tiny, fast-moving details of the atmosphere, such as small eddies and rapid fluctuations, leaving only the larger, slower patterns.

By removing these small, fast details, the training data became much more reversible. The models learned to navigate this simplified version of the atmosphere where the butterfly effect did not exist. In the real atmosphere, tiny errors in the small scales grow quickly and spread to the large scales, ruining a forecast. But in the smoothed-out data the models learned from, those small scales were gone, so the errors did not have a place to hide or grow. This is why the models were so accurate at predicting the future: they were not fighting the rapid error growth that plagues traditional models because the data they learned from had already filtered it out.

The researchers confirmed this by testing the models with different levels of detail. When they trained the models on data that included more of the small, fast scales, the models began to behave more like the real atmosphere. The butterfly effect reappeared, with errors growing rapidly as the initial conditions changed. However, this return to physical realism came at a cost: the models became less accurate at predicting the weather. The very feature that made them good at forecasting—their ability to ignore the chaotic small scales—was also what made them fail to capture the true physics of the atmosphere.

The study also showed that this ability to predict the past was a direct result of the same smoothing process. Because the models learned a version of the atmosphere where the small, chaotic details were missing, they could step backward in time without the errors exploding. In the real world, stepping backward would cause the simulation to blow up instantly, but in the simplified world of the training data, the path remained clear. This suggests that the models are not learning the deep, causal laws of the atmosphere in the way a physicist would. Instead, they are learning a statistical map of the large-scale patterns, effectively treating the weather like a video that can be played forward or backward with equal ease, as long as the fine grain of the film is blurred.

These findings have important implications for how we use and trust artificial intelligence in weather forecasting. The models are incredibly useful tools for short-term predictions because they have learned to ignore the chaotic noise that usually ruins forecasts. But this success comes with a hidden limitation: they may struggle with extreme, unprecedented weather events that depend on those very small-scale details. If a storm is driven by a specific interaction of tiny atmospheric features that the model has never seen because they were smoothed out in the training data, the model might fail to predict it.

The researchers also explored whether these models could be improved by forcing them to learn the missing physics. They tested various approaches, including training the models on higher-resolution data and using different mathematical structures. While these changes made the models behave more like real physical systems, they consistently reduced the models' forecasting skill. This creates a difficult trade-off: a model that is more physically realistic is less accurate at predicting the weather we care about. The study concludes that the "coarse-graining" of the training data is not just a technical detail but a central design choice. It is the reason these models are so successful, but it is also the reason they miss the butterfly effect and the true arrow of time.

Ultimately, the paper suggests that we should view these artificial intelligence models not as perfect simulations of the atmosphere, but as highly skilled pattern recognizers that work within the limits of the data they are given. They have found a way to predict the future by ignoring the chaos that makes the future uncertain. This is a powerful capability, but it is not the same as understanding the physics. As we rely more on these tools for critical decisions, from agriculture to disaster planning, it is vital to remember that their accuracy is built on a simplified version of reality. They are excellent at telling us what will likely happen next, but they may not be able to tell us why, or what happens when the rules of the simplified world no longer apply.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →