← Latest papers
📊 statistics

An Explicit Link between Extreme Value Theory and Compositional Data Analysis

This paper establishes an explicit algebraic link between Extreme Value Theory and Compositional Data Analysis by demonstrating how their respective covariance and dependence representations are connected through simple transformations, thereby enabling the transfer of statistical methods such as graphical models and dimensionality reduction techniques between the two fields.

Original authors: Manuel Hentschel, Sebastian Engelke

Published 2026-07-13
📖 6 min read🧠 Deep dive

Original authors: Manuel Hentschel, Sebastian Engelke

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine two different groups of scientists trying to solve the same puzzle, but they are speaking completely different languages. On one side, you have Extreme Value Theory experts, who study rare, massive events like once-in-a-century floods or sudden stock market crashes. On the other side, you have Compositional Data Analysis experts, who study things where only the proportions matter, like the chemical mix in a rock or the percentage of different species in a forest.

For a long time, these two groups were working in parallel, unaware that they were actually solving the exact same mathematical riddle. This paper, written by Manuel Hentschel and Sebastian Engelke, acts as a universal translator, revealing a hidden bridge between these two worlds.

The Core Discovery: It's All About Ratios, Not Sizes

Here is the big idea: Both fields are obsessed with relative information.

In the world of extreme events, scientists realized that when a massive storm hits, the total size of the storm is less important than the shape of the damage. Did the wind blow harder in the north or the south? The math separates the "radial size" (how big the event is) from the "relative profile" (how the parts compare to each other).

In the world of compositions, scientists realized that if you double the amount of every ingredient in a cake recipe, the cake tastes exactly the same. The absolute amounts don't matter; only the ratios between the flour, sugar, and eggs do.

The paper proves that the mathematical tools used to describe these "shapes" and "ratios" are actually identical. It's like discovering that the map used by a sailor navigating a storm is drawn on the exact same grid as the map used by a baker measuring ingredients.

The Magic Toolkit: Projections and Inverses

The authors didn't just say, "Hey, these look similar." They built a rigorous algebraic framework to show exactly how the tools match up. They use a few fancy-sounding mathematical moves, which we can think of as:

  1. Oblique Projections: Imagine shining a light on a 3D object to cast a shadow on a wall. In this paper, they show how to "project" complex data onto a flat surface to see the relative relationships clearly, ignoring the overall size.
  2. The Variogram Map: This is a special calculator that turns a list of variances (how much things wiggle) into a map of differences. The paper shows that the "extremal variogram" used for floods is the same as the "compositional variation array" used for chemical mixes.
  3. Moore–Penrose Inverses: Think of this as a "reverse button" for math problems that usually can't be reversed because they have too many variables. The paper proves that if you hit this reverse button on the compositional data, you get the exact same result as hitting it on the extreme value data.

What This Means for Real Life (The Transfers)

Because the math is the same, the authors show we can steal tricks from one field and use them in the other.

1. From Extremes to Compositions: Finding Hidden Connections
In extreme value theory, researchers have developed a way to draw "graphs" that show which parts of a system are connected during a crisis. For example, if a flood hits one river, does it automatically hit the next one? The paper introduces a new method called intrinsic logistic-normal graphical models for compositional data.

  • The Test: They tried this on a real dataset of geochemical measurements (the "gemas" dataset). They found that their new method could spot connections between elements just like the extreme value methods do. It's not a magic bullet that solves everything instantly, but in their simulations, it successfully identified the correct connections in a 10-dimensional model, catching 25 out of 27 true links.

2. From Compositions to Extremes: Seeing the Big Picture
Compositional data experts have been using clever ways to shrink huge datasets down to a manageable size (dimensionality reduction) for decades. The paper shows that extreme value researchers can use these same tricks.

  • The Test: They applied a technique called Log-Ratio Analysis to the "Danube" river dataset, which tracks water flow at 38 different stations. By treating the river flow data like a composition, they created a visual map (a biplot) that clearly showed how the river stations were connected. They even tested a "weighted" version, giving more importance to stations with higher water flow, which made the main river path stand out even more clearly.

3. Smoothing the Thresholds
Usually, to study extreme events, scientists have to draw a hard line (a threshold) and say, "Anything below this is normal, anything above is extreme." This is a bit arbitrary.

  • The Suggestion: Borrowing from weighted analysis in compositions, the authors suggest a weighted variogram estimator. Instead of a hard cut-off, they give "more extreme" observations higher weights. In their simulations with 500 datasets, this weighted approach slightly reduced the error (Mean Squared Error) compared to the standard unweighted method, suggesting it might be a better way to handle the messy reality of real-world data.

What They Don't Claim

It's important to know what this paper doesn't say.

  • It doesn't claim to solve the "Zero" problem: Both fields struggle with data that has zeros (like a chemical that isn't present, or a river with no flow). The paper explicitly states that they assumed strictly positive data to make the math work. They suggest that handling zeros is a job for future research, not something they solved here.
  • It doesn't claim the methods work for all distributions: The perfect mathematical link they found works specifically for the Hüsler–Reiss model (a specific type of extreme value distribution). For other, more complex distributions, the connection is messier and involves "exponential tilting," which means the simple translation doesn't work perfectly.
  • It's not a "Breakthrough" in the sense of a finished product: The authors are careful to say they are "introducing" and "exploring" these transfers. They show it can be done and provide simulations to back it up, but they aren't claiming that every statistical problem in these fields is now solved.

The Bottom Line

This paper is a masterclass in connecting dots. It takes two seemingly unrelated fields—one studying the loudest storms and the other studying the finest recipes—and proves they are using the same underlying geometry. By building a bridge of algebra, the authors allow statisticians to swap tools, potentially making it easier to understand both the rarest disasters and the most complex mixtures. It's a solid, mathematically proven link that opens the door for future exploration, rather than a final destination.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →