← Latest papers
🧬 biology

Locus-resolved reconstruction of the pig UGT family reveals an isoform-collapse artifact in frozen ESM-2 retrieval

This study reconstructs the pig UGT gene family to reveal that while frozen ESM-2 embeddings improve the ranking of specific protein isoforms, they do not outperform traditional sequence-search methods in recovering additional genes once evaluation is standardized to the locus level, thereby exposing an isoform-resolution artifact in previous benchmarking.

Original authors: Xiaotong Zhao¹, Hua Chang², Zhuoyu Zhao¹, Xun Xiang¹

Published 2026-08-29
📖 5 min read🧠 Deep dive

Original authors: Xiaotong Zhao¹, Hua Chang², Zhuoyu Zhao¹, Xun Xiang¹

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Inside the cells of every living creature, from the smallest mouse to the largest cow, a quiet chemical cleanup crew is always at work. These workers are proteins called UDP glycosyltransferases, or UGTs for short. Their job is to attach a small sugar molecule to various substances, a process that helps the body neutralize toxins, break down hormones, and manage fats. In livestock like pigs, understanding exactly which of these workers exist and how they function is vital for veterinary medicine and food safety, as it helps scientists predict how animals process drugs or environmental chemicals. However, the pig genome is a complex landscape where these worker genes often appear in crowded clusters, with multiple copies sitting right next to each other. This makes it difficult to tell where one gene ends and another begins, or to distinguish between a full, functional worker and a broken fragment. For years, researchers have struggled to create a clear, accurate map of these specific genes in pigs, often getting confused by the sheer number of similar-looking copies.

A team of researchers at Yunnan Agricultural University has now drawn that map with unprecedented clarity, revealing a surprising twist in how we search for these genes. They focused on a specific family of pig proteins and reconstructed a catalog of sixteen distinct locations, or loci, where these genes live. Thirteen of these locations have complete, working blueprints, while three remain slightly incomplete and require further review. In doing so, they uncovered a hidden trap in modern genetic analysis. Recently, scientists have started using powerful artificial intelligence tools, specifically a type of language model called ESM-2, to find proteins. These tools are excellent at spotting similarities in the chemical language of proteins, even when the sequences look quite different. When the researchers first tested this AI tool, it seemed to outperform traditional search methods, ranking the correct proteins higher and faster. It appeared to be a breakthrough.

However, the researchers discovered that this apparent advantage was an illusion caused by how the data was counted. In biology, a single gene can produce multiple versions of a protein, known as isoforms, much like a single recipe can be written with slight variations. The AI tool was very good at picking the best version of a protein from a specific gene. But when the researchers collapsed all these different versions down to count just the genes themselves, the AI's lead vanished. Once they compared the tools based on the actual number of genes found rather than the number of protein versions, the AI tool performed exactly the same as the older, standard search methods. The AI had not found any new genes; it had simply organized the existing ones better. This finding is crucial because it warns scientists that ranking a specific protein version highly does not mean a new gene has been discovered. To find new genes, one must look at the gene itself, not just its many protein variations.

The study then turned its attention to a particularly messy cluster of these genes on chromosome 8, a region where six similar genes sit side by side. This area is a hotspot for genetic variation and duplication, making it a perfect test case. The researchers combined evidence from the pig's DNA structure, the evolutionary history shared with cattle, and the actual activity of the genes in different tissues. They found that while the genes in this cluster are closely related, they are not identical copies. By using a method that looks for unique short sequences in RNA—the molecule that carries instructions from DNA to build proteins—they could tell which of these genes were actually being used by the pig's body. The results showed that two of the genes in this cluster, named SsUGT2A-like-03 and SsUGT2A-like-04, were clearly active and supported by strong evidence. A third gene, SsUGT2A-like-02, was much harder to detect, with very little evidence that it was being transcribed into RNA, suggesting it might be less active or perhaps a remnant of an older duplication.

The researchers also looked at the evolutionary history of these genes to see if they were constantly changing to adapt to new challenges or if they were staying the same. They found that the genes were under strong pressure to remain unchanged, a sign that they are performing essential, conserved functions rather than rapidly evolving new ones. The three-dimensional shapes predicted for the proteins made by these genes were nearly identical, further confirming that they likely do the same job. This detailed work provides a stable, reliable reference for future studies. Instead of a confusing list of names and overlapping records, scientists now have a clear catalog of sixteen specific locations with verified coordinates and evidence levels. This clarity allows for better design of experiments, more accurate comparisons between species, and a deeper understanding of how pigs handle the chemicals they encounter. The study ultimately teaches a valuable lesson about the tools we use: while artificial intelligence can help us sort through complex data, we must always be careful to ask the right question. If we want to know how many genes exist, we must count the genes, not just the many faces they wear.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →