Tree-aware conditional language modeling recovers mutational patterns of viral evolution
The paper introduces evoPLM-Tree, a tree-aware conditional language model that integrates phylogenetic context to accurately predict viral mutational patterns and lineage-specific evolutionary trajectories, demonstrating strong alignment with experimental functional constraints.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
Viruses are masters of change. They survive by constantly rewriting their own genetic instructions, swapping out parts of their proteins to evade the immune systems of the hosts they infect. This process is not random chaos; it follows a path. A mutation that helps a virus today might be useless or even harmful tomorrow, depending on the other changes that have already happened in that specific family line. Scientists have long tried to predict these changes using computer models that read the language of proteins, treating them like sentences where certain words fit together better than others. However, most of these models look at the final result without paying attention to the journey. They see the destination but ignore the map of how the virus got there, missing the fact that a virus's history shapes what changes it can safely make next.
A team of researchers has developed a new way to look at this problem, one that forces the computer to pay attention to the family tree. They created a system called evoPLM-Tree, which acts like a language model that understands evolution. Instead of just guessing what a protein sequence might look like, this model takes an ancestral virus sequence and a specific family tree as its starting point. It then learns to predict the next step in the lineage, using the history of how the virus has changed over time as a guide. The researchers tested this approach using the spike protein of the SARS-CoV-2 virus, the part that allows the virus to enter human cells. They paired sequences from early versions of the Omicron variant with their ancestors, based on their positions in a phylogenetic tree—a diagram that maps out the evolutionary relationships between different virus strains.
When the team asked the model to generate descendant sequences, it performed with a level of accuracy that standard models could not match. By feeding the model the evolutionary context, the researchers found that the computer relied much more heavily on the actual input information rather than just guessing based on general patterns. The sequences the model created were not just random variations; they accurately reproduced the specific spots on the protein where mutations actually happen in the real world. The model's predictions for how often mutations would appear at different positions matched the observed data with a strong correlation, showing a clear link between what the computer predicted and what nature actually did. This was true for both the entire spike protein and the specific receptor-binding domain, the part of the virus that latches onto human cells.
The study also looked at how well these predictions held up when the virus moved further away from its starting point in time. While the model became less precise at predicting the exact single-letter changes as the evolutionary distance grew, it remained remarkably good at capturing the overall patterns of change across the whole protein. To see if these patterns meant something real, the researchers compared the model's predictions against laboratory experiments that tested which mutations the virus could tolerate without breaking. The mutations that the model assigned a high probability to were found to be the same ones that scientists had already confirmed could survive in a lab setting. The model successfully identified mutations that allowed the virus to maintain its ability to bind to human cells, even though it had never been shown any experimental data during its training.
This work demonstrates that giving a computer model a clear view of a virus's family history allows it to learn the specific rules of that lineage's evolution. By explicitly including the path of ancestry, the model can recover the unique mutational patterns that define how a virus changes over time. The results suggest that this approach provides a reliable framework for understanding protein evolution and could help scientists prioritize which future mutations are most likely to appear, based on the genomic data we already have. It offers a way to see not just what a virus might become, but how it is likely to get there, grounded in the actual history of its ancestors.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.