← Latest papers
📄 evolutionary biology

Ancestral Sequences Cannot be Accurately Reconstructed via Interpolation in a Variational Autoencoder's Latent Space

This study demonstrates that despite the potential of variational autoencoders to model epistatic interactions, their inherent information loss during the decoding process prevents them from accurately reconstructing ancestral sequences, as they are consistently outperformed by standard likelihood-based and parsimony methods.

Original authors: Gorstein, E., Tang, M., Bruzzone, H., Solis-Lemus, C.

Published 2026-09-01
📖 5 min read🧠 Deep dive

Original authors: Gorstein, E., Tang, M., Bruzzone, H., Solis-Lemus, C.

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Deep in the history of life, every living thing carries a molecular record of its ancestors. Scientists have long sought to read this record, a process known as ancestral sequence reconstruction. Imagine trying to guess the exact words of a letter written by a great-grandparent, based only on the copies held by their many great-grandchildren. In biology, these "letters" are the chains of amino acids that make up proteins, the workhorses of every cell. For decades, researchers have used statistical models to fill in the missing words, assuming that each letter in the chain changes independently of the others, like a single word in a sentence being swapped out without affecting the meaning of the whole. This approach has been the standard tool for understanding how proteins evolved over millions of years.

However, biology is rarely that simple. In reality, the letters in a protein chain often depend on one another; changing one amino acid might only be safe if another specific change happens nearby. This complex interdependence, known as epistasis, is difficult to capture with traditional statistical tools. Recently, a new generation of artificial intelligence tools, specifically a type of deep learning model called a variational autoencoder, has offered a different path. These models are excellent at finding hidden patterns in vast amounts of data. They can compress a long protein sequence into a compact, mathematical summary, or "embedding," and then try to rebuild the original sequence from that summary. Because these models learn directly from the data without being told to ignore connections between letters, many scientists hoped they could solve the problem of epistasis and reconstruct ancient proteins with unprecedented accuracy.

A team of researchers set out to test this hope directly. They built a pipeline that used these AI models to guess the sequences of ancient proteins. The idea was straightforward: take the modern proteins we have today, compress them into these mathematical summaries, and then use the known family tree of the proteins to interpolate, or guess, what the summaries of the ancient ancestors looked like. Once they had these guessed summaries, they would feed them back into the AI to generate the actual amino acid sequences. If this worked, it would mean that deep learning could bypass the limitations of traditional statistics and reveal the true history of proteins, even when those histories were shaped by complex interactions between different parts of the molecule.

The researchers put this idea to the test using computer simulations. They created thousands of fake protein families that evolved over time, first using the simple, independent rules that traditional models assume, and then using complex rules that included the interdependent changes found in nature. They then tried to reconstruct the ancestors of these fake families using both the new AI method and the established statistical methods. The results were clear and consistent. In every scenario, whether the proteins evolved simply or with complex dependencies, the traditional statistical methods outperformed the AI approach. The deep learning model failed to produce accurate ancestral sequences, even in the very situations where it was supposed to shine.

The study went deeper to understand why the AI failed. It turned out the problem was not that the AI missed the family tree structure. When the researchers looked at the mathematical summaries, they found that the AI had indeed learned to organize the proteins in a way that reflected their evolutionary relationships. The failure lay in the reconstruction step itself. The AI model, in its effort to compress the data into a small summary, inevitably lost some of the fine details of the original sequence. When the researchers tried to rebuild the ancient proteins from these summaries, the information was too blurry. The model could not generate a sequence with enough precision to match the true history. Even when the researchers increased the size of the mathematical summary to try to keep more information, the gains were small, and the model began to simply memorize the modern proteins rather than learning the general rules of evolution.

The researchers also tested a more advanced version of the AI that was explicitly taught the family tree during its training. They hoped this extra guidance would help, but it did not improve the results. The fundamental bottleneck remained the same: the process of compressing the sequence into a summary and then expanding it back out was too lossy for the delicate task of guessing the past. The study concludes that while these AI embeddings are powerful tools for visualizing protein families and predicting how modern proteins might behave, they are not yet ready to replace the established statistical methods for reconstructing the past. The hope that deep learning could effortlessly solve the problem of complex evolutionary interactions was, in this specific application, a dead end. The traditional methods, despite their simpler assumptions, remain the more reliable way to read the molecular history of life.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →