PartitionFinder-mAIC: Phylogenetic Partitioning using Marginal Akaike Information Criterion
This paper introduces PartitionFinder-mAIC, a new feature in IQ-TREE 3 that utilizes the marginal Akaike Information Criterion to select more efficient partitioning schemes, resulting in fewer partitions and improved phylogenetic inference accuracy compared to traditional AIC and BIC methods.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
Life on Earth is a vast, branching family tree, and scientists spend much of their time trying to draw the branches correctly. To do this, they look at the genetic code found in plants, animals, and microbes, comparing how these sequences have changed over millions of years. However, not all parts of a genetic sequence change at the same speed or in the same way. Some regions are like sturdy old walls that barely shift, while others are like loose bricks that tumble and rearrange frequently. Because of this variety, researchers cannot treat an entire genetic sequence as a single, uniform block. Instead, they must divide the sequence into smaller groups, or partitions, where each group follows its own set of rules for how it evolves. Finding the right way to group these sections is critical; if the groups are too broad, the analysis misses important details, but if they are too narrow, the model becomes cluttered with unnecessary complexity, leading to a confused picture of history.
For years, a popular tool called PartitionFinder has helped scientists decide how to organize these genetic sections. It uses mathematical scores to merge similar groups together, aiming to find a balance that avoids overcomplicating the story. These scores have traditionally relied on methods that treat the grouping of sections as a fixed fact, assuming the boundaries are set in stone before the final history is calculated. But a new approach suggests that this assumption might be too rigid. A recently introduced method, known as the marginal Akaike information criterion, takes a different view. Instead of locking the groups in place, it averages the likelihood of the data across the models of all partitions, treating the boundaries as fluid possibilities. This shift allows the model to better account for the uncertainty inherent in the data, potentially leading to a clearer view of the tree's main branches.
In a new development, researchers have brought this more flexible method into the widely used IQ-TREE software, creating a tool they call PartitionFinder-mAIC. By integrating this new scoring system into the existing algorithms, the team tested whether it could produce better results than the traditional methods. They ran their new tool against a wide range of simulated data, where the true family tree was already known, as well as real-world DNA and protein sequences. The results showed that the new method consistently chose to use fewer partitions than the older approaches. Rather than splitting the genetic data into many small, overly specific groups, it found that broader groupings were sufficient to explain the evolutionary patterns.
When the researchers applied these findings to the evolutionary history of green plants, the impact was tangible. The partitioning schemes generated by the new method led to more accurate inferences at the most important branches of the plant family tree. This suggests that by allowing the model to consider a wider range of possibilities for how the data is grouped, scientists can avoid the trap of overfitting, where a model memorizes the noise in the data rather than learning the true signal. The tool is now available for other researchers to use in the latest version of the software, offering a way to refine the maps of life's history with a method that acknowledges the fluid nature of genetic evolution.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.