← Latest papers
📊 statistics

Perturbation-based Effect Measures for Compositional Data

This paper proposes a novel framework of "average perturbation effects" to address the limitations of traditional parametric approaches in analyzing high-dimensional, sparse compositional data by using hypothetical perturbations to define interpretable, confounding-adjusted statistical functionals that can be efficiently estimated via semiparametric techniques.

Original authors: Anton Rask Lundborg, Niklas Pfister

Published 2026-08-26
📖 3 min read☕ Coffee break read

Original authors: Anton Rask Lundborg, Niklas Pfister

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In many fields of science, researchers study things that are made of parts that must add up to a whole. A geologist might look at the mix of minerals in a rock, an ecologist might count the different species in a forest, or a sociologist might examine the racial makeup of a city. In all these cases, the data is not a list of independent numbers; it is a composition where increasing one part automatically means decreasing the others. This creates a unique mathematical trap. If you try to measure how one specific part affects an outcome, like how the presence of a certain microbe affects health, standard statistical tools often fail. They treat the parts as if they can change independently, which is impossible when they are locked together by the rule that they must sum to one hundred percent. This can lead to confusing or even wrong conclusions about what is actually driving the results.

To solve this, researchers Anton Rask Lundborg and Niklas Pfister have developed a new way to measure cause and effect in these mixed-up systems. Instead of trying to force the data into old models that ignore the "whole" constraint, they proposed a method based on hypothetical changes, or perturbations. Imagine you want to know what happens if you increase the diversity of a group. In the real world, you cannot just magically add more variety without taking something away from the existing mix. The researchers' approach simulates a specific, realistic change: they ask, "What would happen to the outcome if we shifted the entire composition slightly in a specific direction, like moving it toward a more even balance?" By calculating the average result of these hypothetical shifts across many different starting points, they can isolate the true effect of the change without the confusion caused by the other parts of the mix.

The team tested this idea on both simulated data and real-world examples, including a study of racial diversity in New York schools and an analysis of gut bacteria in thousands of people. In the school data, they looked at how racial diversity relates to student grades. Traditional methods, which simply look at the correlation between a diversity score and grades, suggested a strong positive link. However, when the researchers applied their new perturbation method, which accounts for the complex interplay of the different racial groups, the effect was much smaller and, in the case of math grades, statistically indistinguishable from zero. This showed that the strong link seen by older methods was likely an illusion created by other factors, such as the economic background of the students, rather than a direct result of diversity itself.

Similarly, in the study of the human gut microbiome, the researchers wanted to find which specific bacteria are important for predicting a person's body mass index. Standard tools often struggle with this data because many bacteria are absent in some people, creating a "zero" that breaks traditional math. The new method handles these zeros naturally. It distinguished between the effect of a microbe simply being present or absent versus the effect of its abundance changing. The results showed that the new approach identified different, and likely more accurate, bacteria as important compared to older techniques. The researchers found that their method provides a clearer, more honest picture of how changes in a mixture affect an outcome, offering a reliable tool for scientists who work with data that is bound by the rule of the whole.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →