Empirical substitution models of SARS-CoV-2 protein evolution for phylogenetic inference
This study develops and validates novel empirical substitution models specifically for SARS-CoV-2 proteins, demonstrating that they yield more accurate phylogenetic inferences and biologically realistic evolutionary predictions compared to existing generalist models.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
Imagine you are trying to trace the family history of a very famous, very busy celebrity. To do this, you need a rulebook that tells you how likely it is for one family member to change into another over time. In the world of biology, this rulebook is called a "substitution model." Scientists use these models to build family trees (called phylogenies) for viruses, animals, and plants. The goal is to figure out who is related to whom and how they changed over time. However, most rulebooks out there are "generalist" guides—they work okay for many different types of proteins, kind of like a generic "how to cook" book that tries to cover everything from baking a cake to grilling a steak. The problem is that different proteins have very specific jobs and rules. A protein that acts as a virus's engine (a protease) has different rules than a protein that acts as the virus's key to unlock a cell (the spike protein). If you use a generic rulebook for a specific job, your family tree might end up looking a bit wobbly and inaccurate. This is especially true for SARS-CoV-2, the virus that causes COVID-19, which has been evolving rapidly and changing the world. Until now, scientists didn't have a custom rulebook specifically for this virus's proteins, which meant they were trying to fit a square peg into a round hole.
In this study, the researchers decided to write their own, custom rulebooks specifically for SARS-CoV-2. They looked at thousands of real virus sequences to see exactly how the amino acids (the building blocks of proteins) swapped places in the virus's main protease, its papain-like protease, its spike protein, and the whole virus package. Think of it like watching a massive, real-time dance floor of the virus to see which dancers actually switch partners and which ones stay put. They found that the virus has its own unique dance moves that generic rulebooks simply missed. For instance, they discovered that in the virus's proteases (the enzymes that cut up other proteins), certain swaps happen very rarely because the structure is so tight and fragile, while the spike protein is much more flexible and allows for wilder changes, especially those that help the virus hide from our immune system.
The team then tested their new, custom rulebooks against the old, generic ones. They did this in two ways. First, they checked which rulebook could best explain the real data they collected, like a detective checking which theory fits the clues best. Their custom models won almost every time, showing that they are much better at describing how this specific virus evolves. Second, they ran computer simulations where they let the virus evolve forward in time using their new rules versus the old rules. They checked the "folding stability" of the resulting proteins—imagine checking if a paper crane made from the new rules still holds its shape, while the one made from the old rules falls apart or becomes too stiff. The proteins evolved using the new SARS-CoV-2-specific models held their shape much more realistically, looking just like the real virus proteins found in nature.
The researchers conclude that using these new, virus-specific models gives a much clearer and more accurate picture of how SARS-CoV-2 is changing. They suggest that for the best results, scientists should stop using the "one-size-fits-all" rulebooks and switch to these custom guides. While the study shows these models work better in simulations and fit the data more tightly, the authors note that this is a significant step toward understanding the virus's past and predicting its future moves, helping us stay one step ahead in the ongoing battle against the virus.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.