← Latest papers
🧬 biology

A permutation-null artifact drives a conservation–divergence spectrum's signature trend, and phylogeny dominates a protein language model's per-position coupling signal

This paper demonstrates that the signature negative trend in the conservation–divergence spectrum is largely an artifact of its permutation null model and that protein language models' per-position coupling signals are predominantly driven by phylogeny rather than functional mechanism, establishing a rigorous framework for cross-examining black-box models against independent evolutionary evidence.

Original authors: Huazhang Shen

Published 2026-09-01
📖 5 min read🧠 Deep dive

Original authors: Huazhang Shen

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Proteins are the molecular machines that keep life running, and their shapes and functions are dictated by the precise order of their building blocks, called amino acids. For decades, scientists have tried to read these sequences to find the specific spots that determine a protein's unique job. Two very different approaches have emerged to solve this puzzle. The first relies on comparing thousands of related proteins to see which spots stay the same or change together across evolution; this method is transparent and easy to follow, but it can only see patterns written in single letters. The second approach uses artificial intelligence models trained on millions of sequences to predict how a protein behaves; these models are incredibly powerful and can sense complex, hidden patterns, but they are often called "black boxes" because no one is entirely sure what they have actually learned or if they are simply memorizing family trees rather than understanding how the machine works. The question facing the field is whether these intelligent models are discovering new biological truths or just taking shortcuts based on evolutionary history.

A recent study by independent researcher Huazhang Shen brings these two methods face-to-face to test exactly what they know. The researcher focused on a specific group of proteins called G-protein-coupled receptors, which act as switches on the surface of cells, deciding which internal signal to send when they are activated. The goal was to identify the exact spots on these proteins that determine which signal they choose. The study treated the traditional evolutionary method and the artificial intelligence model as two independent observers examining the same set of proteins. One observer looked for specific amino acid changes that distinguish the different signal types, while the other used a neural network to guess the signal type based on the protein's sequence. The researchers then compared the lists of important spots generated by each observer to see if they agreed.

The comparison revealed a surprising lack of agreement. The two methods pointed to almost entirely different spots as being important, with their rankings barely overlapping. This divergence suggested that the two observers were looking at the problem through completely different lenses. To understand why, the researcher first examined the traditional evolutionary method. This method had previously shown a clear pattern: spots that were highly conserved across evolution tended to be less variable between different signal types. While this looked like a biological law, the study found that this pattern was largely an illusion created by the mathematics of the test itself. By shuffling the data to create a statistical baseline, the researcher showed that the negative trend appeared even in random noise. The real biological signal was not the trend itself, but the small amount of extra variation that remained after subtracting this mathematical baseline. This finding implies that many previous studies claiming to find specificity based on this trend might have been measuring a statistical artifact rather than a biological rule.

The investigation then turned to the artificial intelligence model to see what it was actually learning. When the model was tested without any restrictions, it appeared to know a great deal about which signal a protein would choose. However, when the researcher forced the model to prove it understood the mechanism rather than just recognizing evolutionary cousins, the results changed dramatically. By grouping related proteins together during testing so the model could not rely on family resemblance, the model's performance dropped significantly. The study calculated that roughly three-quarters of the model's apparent knowledge was simply a reflection of evolutionary relatedness, not the actual mechanism of how the protein works. Only about one-quarter of its signal represented genuine functional insight. This provides strong quantitative evidence for the criticism that these powerful models often act as shortcuts for recognizing homology rather than true understanding of biological function.

Despite these limitations, the study found a specific area where the artificial intelligence model held an advantage. When the two methods disagreed, the model's predictions were slightly more likely to align with the physical locations where the protein actually touches its internal partner, based on known 3D structures. However, this alignment was weak and appeared as a general shift in the data rather than a perfect list of correct answers. To ensure the methods were working correctly, the researcher also tested them on a different family of proteins where the biological code is known to be too complex for either method to read. In this case, both methods failed to find the correct spots and disagreed with each other, confirming that the framework could distinguish between a genuine biological limit and a failure of the tools.

The study concludes with a new framework for testing what these black-box models actually know. It demonstrates that before trusting a model's predictions, scientists must strip away the influence of evolutionary history and account for statistical baselines that can mimic real patterns. While the artificial intelligence model does capture some real functional information, it is heavily contaminated by family history, and the traditional evolutionary method is prone to misinterpreting statistical noise as biological law. The research does not declare one method a winner but instead provides a rigorous way to cross-examine them, ensuring that future discoveries about how proteins work are based on genuine biological signals rather than mathematical illusions or evolutionary shortcuts.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →