← Latest papers
🧬 biology

Reconstructing sequence-grammar trajectories enables interpretable and tunable cis-regulatory element design

The paper presents GO-CRE, an interpretable deep learning framework that combines a hybrid Transformer-Mamba2 model with reinforcement learning and sequence-grammar trajectory reconstruction to design and experimentally validate diverse, cell-type-specific cis-regulatory elements with enhanced activity and specificity.

Original authors: Mingqian Ma, Wanjuan Bu, Guoqing Liu, Yuxuan Liu, Sizhen Liu, Zhen Zhao, Shijie Yao, Qingru Hua, Yujie Zhang, Cuiting Zhong, Haitao Huang, Pan Deng, Peiran Jin, Qijin Yin, Chuan Cao, Haiguang Liu, Mo
Published 2026-09-23
📖 6 min read🧠 Deep dive

Original authors: Mingqian Ma, Wanjuan Bu, Guoqing Liu, Yuxuan Liu, Sizhen Liu, Zhen Zhao, Shijie Yao, Qingru Hua, Yujie Zhang, Cuiting Zhong, Haitao Huang, Pan Deng, Peiran Jin, Qijin Yin, Chuan Cao, Haiguang Liu, Mo Xu, Yuan He, Tao Qin, Zeyu Chen

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Inside every cell, a vast library of instructions dictates how life functions, but the most critical parts of this library are not the genes themselves. Instead, the true conductors of cellular life are short, switch-like segments of DNA known as cis-regulatory elements. These segments act as dimmer switches and on-off buttons, telling specific genes when to turn on, how loudly to sing, and in which types of cells to operate. Without them, a liver cell could not know to filter toxins, and a blood cell could not know to carry oxygen. For decades, scientists have been able to read these switches, but designing new ones from scratch to control cells for therapies or research has remained a profound challenge. It is like trying to write a new language where the grammar is invisible and the rules change depending on the neighborhood the word lives in.

A team of researchers has now developed a new way to design these genetic switches, turning a process that was once a black box into a transparent, steerable journey. By building a system that watches how DNA sequences evolve as they are being designed, they can see exactly how the "grammar" of life emerges. This approach allowed them to create synthetic switches that work with high precision in specific human cell types, offering a clearer path toward engineering the genetic circuits that control our health.

The researchers, led by Mingqian Ma and colleagues, created a framework they call GO-CRE, which stands for Guided Optimization of Cis-Regulatory Elements. To understand how it works, imagine a computer program trying to write a sentence that will make a specific cell type react in a desired way. In the past, scientists would let an artificial intelligence generate thousands of DNA sequences, test them, and hope the best ones worked. This was often a trial-and-error process where the computer might find a solution that worked by accident, using simple, repetitive patterns that fooled the test but didn't truly understand the biology. The new system changes this by treating the design process not as a race to a finish line, but as a map of a journey.

The core of their method relies on a powerful computer model called HybriDNA. This model was trained on a massive collection of DNA from hundreds of species, learning the subtle patterns that govern how genetic information is stored and read. The researchers then used this model to generate new DNA sequences, but they did not stop there. They added a layer of observation that tracks every small change the computer makes as it tries to improve a sequence. They broke these changes down into tiny building blocks, looking at how the frequency of specific letter combinations and the presence of known biological tags shifted over time. This allowed them to reconstruct the "trajectory," or the path, the sequence took as it evolved from a random string of DNA into a functional switch.

When the team watched these paths unfold in a liver cell type known as HepG2, they discovered something unexpected. The sequences were getting stuck in a trap. The computer was finding a shortcut by creating long, repetitive stretches of a single letter, specifically a string of Gs, which the model mistakenly thought was a good solution. This was a low-quality answer that satisfied the computer's scoring system but failed to create a real, working biological switch. Because the researchers could see this happening in real-time, they were able to intervene. They adjusted the rules of the game, adding a penalty for these repetitive strings. This simple correction redirected the computer's path, steering it away from the trap and toward a more complex, biologically meaningful solution.

This ability to diagnose and fix the design process in the middle of the journey revealed that successful creation of these switches happens in three distinct phases. First, the computer explores widely, trying many different combinations. Next, it commits to a specific direction, locking onto a set of biological features that belong to the target cell type. Finally, it fine-tunes the result, polishing the sequence while keeping its core identity. In the liver cells, the system learned to assemble a specific team of regulatory tags in a precise order, creating a compact and efficient switch. In blood cells, known as K562, it assembled a different team of tags suited for that environment.

To prove that these computer-generated switches actually worked, the researchers tested them in the lab. They inserted the new DNA sequences into living liver and blood cells using a harmless virus. They then measured how well these new switches turned on a glowing reporter gene. The results were striking. The switches designed for liver cells lit up brightly in liver cells but remained dark in blood cells, and vice versa. This confirmed that the computer had successfully learned to write cell-type-specific instructions. Furthermore, the new liver switches were, on average, more active than the natural switches found in the human genome, suggesting that the computer had found a more efficient way to organize the genetic code.

The study also highlighted that not every cell type was equally easy to design for. While the system succeeded brilliantly in liver and blood cells, it struggled to create specific switches for a neural cell type. In this case, the computer's path remained scattered and failed to settle on a clear pattern. This failure was just as informative as the success; it showed that the system is limited by the quality and quantity of data available for that specific cell type. It revealed that for the computer to learn the grammar of a cell, it needs a rich library of examples to study, and without enough data, the design process cannot find a stable path.

By making the design process visible and adjustable, this work transforms how scientists approach genetic engineering. Instead of hoping for a lucky result, they can now watch the sequence evolve, spot when it is going off track, and guide it back to a productive path. The resulting switches are not just random sequences that happen to work; they are the product of a guided evolution that converges on compact, efficient, and highly specific biological programs. This approach offers a powerful new tool for creating the precise genetic controls needed for future therapies, turning the abstract task of writing DNA into a tangible, understandable, and controllable process.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →