Identification and Inference for Causal Effects in Extremes under General Conditions
This paper develops identification and inference methods for causal effects in extreme events by analyzing the Causal Tail Coefficient within linear structural models, demonstrating how heterogeneous tail indices and heavy-tailed confounders influence causal structure identification and proposing adjusted estimators and tests to recover causal relations that average-based methods often miss.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the world of data science, researchers have long relied on a simple principle: to understand how one thing affects another, look at the average. If you want to know if rain causes traffic jams, you might check how much slower cars move on a typical rainy day compared to a sunny one. This approach works well for everyday fluctuations, but it often fails when the stakes are highest. The most dangerous events in finance, climate, and infrastructure—market crashes, catastrophic floods, or systemic grid failures—do not happen on average. They happen in the extremes, in the rare, violent outliers where the usual rules of cause and effect can break down or hide entirely. For decades, statisticians struggled to untangle these extreme relationships because the tools they used were designed for the middle of the data, not the fringes. They could tell you how variables moved together in general, but they could not reliably say which one was pushing the other when everything was falling apart.
This is the problem tackled by a new study from researchers at the Karlsruhe Institute of Technology and Heidelberg Institute for Theoretical Studies. They developed a method to trace the path of extreme events through a system, asking a specific question: when a massive shock hits one part of a network, does it trigger a massive shock in another? To do this, they focused on a concept called the "causal tail coefficient." Imagine two variables, like rainfall and river levels. If a huge rainstorm consistently leads to a massive flood, there is a causal link. But if both are simply reacting to a third, hidden factor—like a massive weather system affecting the whole region—the link is an illusion. The researchers found that by looking specifically at the behavior of the most extreme values, they could distinguish between a true cause-and-effect chain and a coincidence caused by a hidden influence. Their work is particularly powerful because it handles real-world messiness: it works even when the two variables behave differently at the extremes, such as when one variable has a "heavier" tail, meaning it produces more frequent or more severe outliers than the other.
The researchers tested their theory using a framework where variables are connected in a specific direction, like a chain of dominoes. They discovered that the way these variables behave at the very edge of the distribution holds a secret key to identifying the direction of causality. If one variable has a "heavier" tail than the other—meaning it experiences more extreme spikes—their method can often tell you which one is the driver and which one is the follower. In cases where the tails are similar, the method still works, but it requires looking for a specific pattern in how the extremes align. Crucially, they also addressed the problem of hidden confounders. If a third, unobserved factor is causing both variables to spike, standard methods often mistake this for a direct cause. The team showed that if this hidden factor is also extreme (heavy-tailed), it can completely mask the true relationship. However, they devised a way to correct for this if researchers have a proxy—a measurable stand-in—for that hidden factor. By adjusting their calculations to account for this proxy, they could peel back the layer of confusion and reveal the true causal arrow, even in the presence of heavy-tailed noise.
To prove their method worked, the team ran thousands of computer simulations, creating artificial worlds with different types of extreme behaviors and hidden influences. They found that their new tests could correctly identify causal directions in situations where older, standard methods failed completely. For instance, in scenarios where a causal link existed only in the extreme tails of the data, traditional tools saw nothing, while the new method spotted the connection clearly. They also tested how well the method performed when the "tails" of the data were different, confirming that these differences actually helped them pinpoint the cause rather than confusing the issue. The simulations showed that the method remains reliable even when the data is noisy or when the sample size is not massive, provided the researchers choose the right threshold for what counts as an "extreme" event.
The researchers then took their tools out of the simulation lab and applied them to real-world data, starting with the Swiss railway system. They wanted to know if extreme precipitation caused extreme train delays. While it seems obvious that rain delays trains, standard statistical models looking at average data often struggle to prove this because delays happen for many reasons. By focusing only on the most extreme rain events and the most severe delays, the researchers found a clear, one-way causal link: heavy rain directly caused significant delays on the Zurich-to-Bern line. The method also revealed that the impact of rain accumulated over three hours was even more significant than the rain falling at a single moment, a nuance that average-based models missed.
Next, they turned to the rivers of Germany, examining the relationship between daily rainfall and river discharge on the Danube and the Main rivers. Here, the challenge was more complex. The data suggested that while rain caused the rivers to swell, there was also a hidden factor at play, likely related to the upstream catchment areas. The researchers' method detected this hidden influence, which they confirmed by checking the "tail indices" of the data. Once they adjusted their analysis to account for the upstream water flow as a proxy for the hidden factor, the true causal link between rain and river levels became sharp and undeniable. They also explored how long it took for rain to affect the river, finding that while the connection was strong after one day, the causal signal in the extremes faded much faster than the average signal did, disappearing after just a few days. This provided a more precise picture of flood dynamics than previously available.
Finally, the team looked at the volatile relationship between the S&P 500 stock index and Bitcoin. Financial markets are notorious for their extreme swings, and the question of whether stock market crashes drive cryptocurrency crashes, or vice versa, has been a subject of intense debate. Previous studies using average data had produced conflicting results. The researchers applied their extreme-event method to the left tail of the data—the side representing massive losses. They found that when the stock market suffered extreme negative returns, it triggered extreme negative returns in Bitcoin, but only after they accounted for broader market volatility as a hidden confounder. Without this adjustment, the signal was muddled, but with it, a clear causal path emerged from the stock market to the cryptocurrency. This suggests that during times of severe market stress, the traditional stock market acts as a leading indicator for the crypto market, a finding that could be vital for investors and regulators trying to understand systemic risk.
The study concludes that looking at the extremes is not just a niche exercise but a necessary tool for understanding how the world works when it is under the most pressure. By developing a way to separate true causal chains from coincidental correlations in the tails of distributions, the researchers have provided a new lens for economists, climatologists, and risk managers. Their work demonstrates that the most extreme events often carry the clearest signals of cause and effect, provided one knows how to read them. The method does not require perfect data or simple systems; it thrives in the messy, heavy-tailed reality of the natural and financial worlds, offering a way to see the true structure of causality when everything else is in chaos.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.