Group invariance of -divergences and the Fisher--Rao distance
This paper demonstrates that -divergences and the Fisher--Rao distance are invariant under group actions in transformation models, thereby reducing their dependence on parameter pairs to maximal invariants such as double cosets, particularly within multidimensional location-scale families.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a chef trying to compare two different soups. One is a small bowl of spicy tomato soup, and the other is a giant pot of mild tomato soup.
If you just look at the bowls, they look totally different. But if you know the "recipe" (the statistical model), you realize that the only real differences are how much you stretched the pot (scale) and where you moved it on the counter (location). If you ignore the size of the pot and the counter's position, the soups are essentially the same "flavor profile."
This paper is about finding a mathematical way to ignore those irrelevant differences (like pot size and counter position) so we can measure the true difference between two distributions (like two soups) fairly.
Here is a breakdown of the paper's main ideas using simple analogies:
1. The Problem: Too Many Ways to Describe the Same Thing
In statistics, we often describe data using parameters (like a mean and a variance). But nature has "symmetries."
- Translation (Location): If you move a whole dataset 5 steps to the right, the relationship between two points in that data shouldn't change.
- Stretching (Scale): If you zoom in or out on a dataset, the shape of the relationship shouldn't change.
The authors ask: If we have a mathematical tool to measure the "distance" or "difference" between two distributions, does that tool respect these symmetries? In other words, if we stretch or move both distributions by the exact same amount, should the distance between them stay the same?
2. The Big Discovery: The "Group" Magic
The paper proves that a huge family of these distance-measuring tools, called f-divergences (which includes famous ones like Kullback-Leibler divergence), automatically respect these symmetries.
The Analogy: Imagine you have a rubber sheet with two dots drawn on it.
- If you stretch the whole sheet (scale) or slide it across the table (location), the distance between the two dots on the sheet doesn't change.
- The paper shows that for these specific statistical tools, the "distance" is like the measurement on the rubber sheet, not the measurement on the table. It is invariant.
3. The Solution: The "Maximal Invariant" (The Ultimate ID Card)
Just knowing that the distance doesn't change when we stretch or move things is good, but it's not enough. We want to know: What exactly determines the distance?
The authors introduce a concept called a Maximal Invariant. Think of this as an "Ultimate ID Card" for a pair of distributions.
- If you have two pairs of distributions, and they have the same "Ultimate ID Card," then the distance between them is identical.
- If they have different ID cards, the distance might be different.
This ID card strips away all the "noise" (the stretching and moving) and leaves only the relative information.
- Relative Location: How far apart are the centers?
- Relative Scale: How much bigger is one than the other?
4. The "Double Coset": A Geometric Shortcut
When the authors look at complex, multi-dimensional data (like 3D shapes instead of just 1D lines), they find that this "Ultimate ID Card" has a specific geometric shape called a Double Coset.
The Analogy: Imagine you are trying to describe the position of two people in a room relative to each other, but the room is rotating and the people are walking around.
- Instead of tracking their absolute coordinates (which change constantly), you track their relative pose.
- The "Double Coset" is the mathematical name for that relative pose. It tells you everything you need to know about how the two distributions are positioned relative to one another, ignoring the fact that the whole universe might be spinning.
5. Applying it to Real Models (Location-Scale Families)
The paper applies this to Location-Scale families (distributions defined by a center and a spread).
- They show that for these models, the "Ultimate ID Card" is made of two things:
- Singular Values: These are like the "stretch factors" in different directions.
- Block Norms: These measure the "relative shift" of the centers, but only in the directions where the stretching happened.
This means you don't need to calculate a complex formula for every possible rotation or shift. You just need to calculate these specific numbers (singular values and block norms), and that is all the information you need to know the distance.
6. The Fisher-Rao Distance (The "Curved" Distance)
Finally, the paper looks at the Fisher-Rao distance, which is a special kind of distance used in information geometry. It's like measuring the distance between two points on a curved surface (a geodesic) rather than a flat map.
The authors prove that this curved distance also follows the same rules. Even though the math is harder, the result is the same: the distance between two distributions depends only on their relative position and scale (the "Ultimate ID Card"), not on where they are in the universe.
Summary
The paper is a guide to simplifying complex comparisons.
- Symmetry exists: Moving or stretching data doesn't change the intrinsic difference between two distributions.
- Tools respect symmetry: Common statistical distance tools (f-divergences) naturally ignore these moves.
- The "ID Card": We can replace complex parameters with a simpler set of numbers (singular values and relative shifts) that capture the only thing that matters: the relative relationship between the two distributions.
- Universal Application: This works for both standard distance measures and the more complex, curved Fisher-Rao distance.
In short: To compare two things fairly, ignore where they are and how big they are; look only at how they relate to each other. The paper gives the mathematical recipe for doing exactly that.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.