← Latest papers
🧠 neurology

Evaluation of multiple sclerosis risk loci reveals that genomic oracles fail to generalise sequence effects across host contexts

This study demonstrates that deep-learning genomic oracles fail to generalize sequence effects across different host contexts, as their predictions for regulatory element activity are highly dependent on the specific genomic location and cannot be reliably predicted by sequence features alone.

Original authors: Selahattin Alperen Uysal

Published 2026-09-21✓ Author reviewed ⓘ
📖 5 min read🧠 Deep dive

Original authors: Selahattin Alperen Uysal

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the vast landscape of the human genome, certain stretches of DNA act as switches, turning genes on or off to control how cells function. Scientists have long sought to understand the precise code that tells these switches where to work and when to activate. In recent years, powerful computer programs have emerged to help decode this language. These programs, often called "genomic oracles," can look at a string of DNA letters and predict how it will behave inside a living cell, estimating whether it will open up to allow gene activity or stay closed. Because these digital tools are fast and cheap compared to laboratory experiments, researchers increasingly use them to design new DNA sequences, hoping to create custom switches that work exactly as intended. The underlying hope is that if a computer program says a specific DNA sequence is a good switch in one part of the genome, it will remain a good switch if moved to a different location.

A new study challenges this fundamental assumption by testing whether these digital predictions hold up when the same DNA sequence is placed in different neighborhoods of the genome. The researchers focused on multiple sclerosis, a disease where the immune system attacks the nervous system, and looked at the specific genetic regions known to increase the risk of developing the condition. They used an artificial intelligence model to generate thousands of candidate DNA sequences that looked like they could function as regulatory switches in immune cells. They then used a genomic oracle to score these candidates, selecting the ones the computer predicted would be the most active. The team then performed a rigorous test: they took the exact same selected sequences and placed them into twenty-four different locations within the genome to see if the computer's prediction remained consistent.

The results revealed a startling lack of consistency. While the computer program could rank sequences well within a single location, its predictions failed to generalize when the sequences were moved. The effect of a specific change to the DNA sequence was not a fixed property of the sequence itself; instead, it depended entirely on where in the genome the sequence was sitting. In some locations, a specific edit to the DNA made the sequence more active, while in other locations, the exact same edit made it less active or had no effect at all. The researchers found that the behavior of the sequence was so dependent on its surroundings that the variation between different locations was massive, with the effect reversing direction in nearly half of the tested sites.

To understand what the computer was actually "seeing," the team performed direct interventions on the DNA sequences. They tested whether the computer's preference was driven by the presence of specific binding sites for proteins known to control genes, or by the density of these sites. The answer was no; removing these sites or changing their number did not change the computer's score in a predictable way. Instead, the factor that truly mattered was the content of a specific chemical structure called CpG, which involves a pairing of two DNA letters. However, even this factor did not act alone. Increasing the amount of this structure in the DNA improved the score in some genomic locations but hurt the score in others. The direction of the effect was determined by the host location, not by the edit itself.

The study also tested whether a second, independently trained computer program would produce the same results. This new model, built with a different architecture and trained on different data, confirmed the main finding: the effect of the DNA edit was highly unstable and varied wildly depending on the genomic location. However, the two programs did not agree on exactly which locations would show a positive or negative effect, suggesting that while the instability is a real phenomenon, the specific details of how a model interprets a location are unique to that model.

Perhaps most critically, the researchers checked if the sequences selected by the computer actually worked better in a real-world experiment. They compared the computer's scores against measured activity from a large-scale laboratory assay involving thousands of DNA elements. The computer's selection criteria failed to predict which sequences were truly active in the cells. In fact, the sequences the computer rated as the best performers often showed the opposite behavior in the lab, with the most specific and active elements receiving the lowest scores from the oracle. This disconnect suggests that the specific goal the computer was optimizing for does not reflect the actual biological activity of the DNA.

The researchers also discovered that a much simpler mathematical method, based on counting patterns of letters in the DNA, could generate sequences just as well as the complex, deep-learning artificial intelligence models. This simple method, which requires no powerful computers and can be fitted in seconds, matched the performance of the advanced neural networks on every measure tested. This finding implies that the complexity of the deep-learning models was not necessary to capture the essential patterns of the DNA sequences in this context.

Ultimately, the study concludes that a sequence-level effect learned or measured by a genomic oracle is not a property of the sequence alone. It is a property of the sequence combined with the specific genomic context in which it sits. This dependence on location is not predictable from simple, cheap features of the surrounding DNA. For scientists designing new genetic elements, this means that a design chosen and validated in a computer simulation at one location carries no guarantee about how it will behave once moved or inserted elsewhere. The findings serve as a caution that these powerful digital tools, while useful, cannot yet be trusted to predict the behavior of designed sequences across the diverse landscape of the human genome without direct experimental verification at multiple locations.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →