Quantifying internal variability in large-ensemble climate data with the Wasserstein distance
This paper proposes a robust, parameter-free metric based on the 1-Wasserstein distance to quantify the relative magnitude of internal variability in large-ensemble climate data, demonstrating its superior stability and clarity compared to existing methods across various distribution shapes and forcing scenarios.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Earth's climate is never still. It breathes with a rhythm of its own, shifting from year to year due to the chaotic, swirling interactions of the atmosphere and oceans. This natural churning is called internal variability. At the same time, the planet is being pushed by external forces, primarily the heat-trapping gases humans release, which drive a long-term trend of warming. To understand the future, scientists must separate these two signals: the steady hand of human influence and the wild, unpredictable hand of nature. If the natural noise is too loud, it can hide the signal of change, making it difficult to know exactly how much the climate will shift or how certain we can be about those predictions.
For decades, researchers have used statistical tools to measure how much of the climate's movement is due to this internal noise versus external forcing. These tools often rely on calculating the average spread of data, a method that works well when the data follows a predictable, bell-shaped pattern. However, many climate variables, such as rainfall, do not follow this neat pattern; they are often skewed, with rare but extreme events that throw off simple averages. When the data is messy, traditional methods can struggle to give a clear picture, sometimes missing the subtle ways in which human influence is beginning to dominate the natural chaos.
In a new study, researchers Yuki Yasuda and Shoichiro Kido from the Japan Agency for Marine-Earth Science and Technology propose a fresh way to measure this balance. They introduce a new metric based on a concept from mathematics called optimal transport, which essentially asks: how much effort does it take to reshape one distribution of data into another? Imagine two piles of sand with different shapes; the cost to move the grains from one shape to the other is a precise measure of how different they are. The researchers apply this idea to climate data by comparing the spread of all individual climate simulations against the spread of their average. This allows them to quantify the relative size of the internal noise without needing to force the data into a specific shape or rely on arbitrary choices about how to group the numbers.
The team tested their new method using synthetic climate data generated from computer models that mimic real-world weather patterns, including those with heavy tails and extreme outliers. They compared their approach against two existing methods: one based on simple variance (spread) and another based on information theory. The results showed that the traditional variance method is highly sensitive to extreme outliers, while the information-based method can be unstable depending on how the data is grouped. In contrast, the new metric provided stable and consistent estimates across different types of data, provided the ensemble of simulations was large enough—specifically, containing about 40 or more members. This threshold offers a practical guideline for scientists designing future climate experiments, suggesting that smaller groups of simulations may not yield reliable results when using this new tool.
The researchers then applied their metric to a massive dataset known as the Community Earth System Model Large Ensemble, which contains 40 simulations of the Earth's climate under two different scenarios: historical conditions from the 20th century and a future scenario with high greenhouse gas emissions. When looking at air temperature, all three methods agreed that the relative contribution of internal noise decreases as the external forcing increases. However, the new metric showed this change most clearly, offering a sharper signal of the human influence on the climate. The difference was even more striking when analyzing total precipitation. Because rainfall data is notoriously irregular and heavy-tailed, the traditional methods failed to detect any significant change in the balance between noise and signal. The new metric, however, successfully revealed that the relative contribution of internal variability is decreasing in specific regions, such as the high latitudes and parts of the tropical oceans, as the climate warms.
This work suggests that the new metric is a powerful, simple tool for analyzing large climate datasets, particularly when dealing with complex, non-standard distributions like precipitation. It does not require complex parameters or assumptions about the shape of the data, making it accessible and robust. While the study confirms the metric's effectiveness in simulations and existing model data, the authors note that further work is needed to fully understand the specific causes of the internal and forced variability it detects. As climate models grow larger and more complex, tools that can cut through the noise without distortion will be essential for distinguishing the steady march of climate change from the natural fluctuations of the Earth system.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.