A large-scale evaluation of tree shape indices reveals potential pitfalls
This study presents a large-scale evaluation of 54 phylogenetic tree shape indices on over 45 million empirical trees using the new Python library "treeshapy," revealing critical pitfalls such as root sensitivity, size dependency, and index redundancy to guide more cautious and effective future tree shape analyses.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
Life on Earth is connected by a vast, branching history, a family tree that stretches back billions of years. Scientists often represent this history as a diagram where lines split to show how species diverged from common ancestors. These diagrams, known as phylogenetic trees, are more than just pictures of who is related to whom; their overall shape tells a story about how evolution happened. Some trees are broad and bushy, suggesting many species evolved quickly at the same time, while others are long and spindly, hinting at a slow, steady march of change. For decades, researchers have tried to measure these shapes using mathematical tools called indices. These numbers are meant to capture the essence of a tree's structure, allowing scientists to compare different groups of life and test theories about how biodiversity grows. However, until now, no one had tested whether these tools actually work reliably on the millions of real-world trees that exist in scientific databases.
A team of researchers recently decided to put these measuring tools to the ultimate test. They gathered more than 45,000,000 real evolutionary trees from two major public collections, a scale of data never before attempted. To handle this massive amount of information, they built a new, open-source software tool designed specifically to calculate 54 different shape indices. Their goal was simple but critical: to see if these numbers hold up when applied to the messy, complex reality of actual biological data, rather than just theoretical models. What they found suggests that many scientists have been drawing conclusions based on measurements that are far more fragile than previously thought.
The study revealed that the shape of a tree is often not a fixed property but depends heavily on where the starting point, or root, is placed. In many evolutionary trees, the exact location of the root is uncertain or debated. The researchers discovered that for 14 of the 54 indices they tested, moving the root even slightly caused the shape value to change dramatically. Only five of the indices remained stable regardless of where the root was placed, while the rest showed a moderate level of sensitivity. This means that if a researcher is unsure about the root of a tree, the shape they calculate might be an artifact of that uncertainty rather than a true reflection of evolutionary history. The values are not just slightly off; they can be fundamentally misleading depending on how the tree is oriented.
Another major finding concerns the size of the tree itself. The researchers observed that almost all of the indices, with the exception of five, are inherently linked to the number of species in the tree. Even when the scientists tried to use standard techniques to normalize the data and remove the effect of size, the correlation remained. This creates a significant problem for comparison: a tree with a hundred species and a tree with a thousand species cannot be fairly compared using these tools, because the numbers will differ simply because one tree is larger, not because their shapes are fundamentally different. It is like trying to compare the density of a small pebble to a large boulder using a ruler; the size difference overwhelms the shape difference.
The researchers also found that many of these 54 indices are not independent of one another. Instead, large groups of them move in lockstep, rising and falling together to describe the same aspect of the tree. This redundancy means that using a long list of indices does not necessarily provide a more complete picture; it often just repeats the same information in different ways. To get a true understanding of a tree's shape, the study suggests that scientists must carefully select a small, diverse set of uncorrelated indices. This approach would capture the different facets of the tree's structure without wasting effort on redundant data.
Ultimately, this large-scale evaluation serves as a necessary guide for the future of evolutionary biology. It does not dismiss the value of tree shape analysis, but it demands a more cautious and rigorous approach. By identifying which tools are stable and which are prone to error, the study provides a clear path forward. Scientists can now choose the right indices for their work, ensuring that their conclusions about the history of life are based on solid ground rather than on measurements that shift with the slightest change in perspective. The researchers have made their new software available to the public, inviting the scientific community to adopt these more careful standards and to interpret the shapes of life's history with greater precision.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.