← Latest papers
📄 bioengineering

RNASeek: A Cross-Phyla Generative Foundation Model for Multipurpose RNA Modeling and Reinforcement Learning-Based Design

RNASeek is a 1.6-billion-parameter generative foundation model that leverages reinforcement learning to unify RNA sequence representation, functional prediction, and de novo design, successfully generating experimentally validated ribozymes and 3' UTRs with enhanced performance.

Original authors: Chen, S., Fernandes, N., Li, W. V., Wang, L., Jia, Z., Xiang, J. S.

Published 2026-09-28
📖 5 min read🧠 Deep dive

Original authors: Chen, S., Fernandes, N., Li, W. V., Wang, L., Jia, Z., Xiang, J. S.

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Inside every living cell, a vast and intricate conversation takes place, not with words, but with molecules. At the heart of this dialogue is RNA, a versatile molecule that acts as both a messenger and a regulator, carrying instructions from the genetic blueprint and deciding how those instructions are executed. For decades, scientists have understood that the specific order of letters in an RNA strand determines its shape and function, much like the arrangement of letters in a sentence determines its meaning. However, predicting exactly how a new sequence of letters will behave has been a formidable challenge. The rules are complex, involving long-range interactions and subtle structural nuances that are difficult to map by hand. Now, researchers have developed a new tool that learns these rules by reading the biological literature of life itself, allowing them to not only predict how RNA will behave but to design entirely new molecules with specific, desired properties.

The team behind this work, led by researchers at the University of California, Riverside and Columbia University, created a massive computer program they call RNASeek. Think of this program as a student that has read every book in a library containing the genetic instructions of over one hundred different species, from plants and animals to fungi and viruses. Unlike previous tools that were designed to solve only one specific puzzle, RNASeek was built to understand the general language of RNA. It was trained on a dataset comprising roughly 150 billion pieces of genetic information, learning to recognize patterns in how different organisms construct their RNA. The program also learned to read natural language instructions, meaning a scientist can simply tell it what they want, such as "create a stable RNA sequence" or "design a molecule that cuts itself," and the model will attempt to generate a solution.

To test if this digital student had truly learned the language, the researchers first asked it to predict the behavior of ribozymes. These are RNA molecules that act as enzymes, capable of cutting themselves or other molecules. The team provided the model with thousands of known ribozyme sequences and their measured cutting speeds. RNASeek successfully learned to predict which sequences would cut quickly and which would be slow, identifying subtle features that contribute to speed, such as the flexibility of certain loops in the molecule's structure. The model performed as well as, or better than, many specialized tools designed specifically for this task, proving that it had grasped the underlying principles of how RNA structure dictates function.

Encouraged by this success, the researchers pushed the model further, asking it to design new RNA sequences from scratch. They focused on a different type of RNA found in the 3' untranslated region, a section of the molecule that helps determine how long the RNA survives inside a cell. A longer lifespan means the cell produces more of the protein the RNA codes for, which is crucial for therapies like mRNA vaccines. The researchers used a technique called reinforcement learning, a method where the computer program plays a game of trial and error. It generated thousands of candidate sequences, and a separate part of the model acted as a judge, scoring them based on how stable they were predicted to be. The program then adjusted its strategy to generate better candidates, gradually learning to produce sequences that were highly stable.

The results were striking. The sequences designed by RNASeek were not just random combinations of letters; they contained specific patterns, such as an abundance of adenine and uracil, that are known to help stabilize RNA. When the researchers tested these computer-generated designs in a laboratory setting, they found that the new sequences performed exceptionally well. Some of the designed molecules were even more stable than the best sequences found in nature or those created by other artificial intelligence tools. Crucially, when the researchers shuffled the letters of these successful designs, scrambling their order while keeping the same ingredients, the stability vanished. This confirmed that the success came from the specific arrangement of the letters, not just the chemical composition, proving the model had learned the true syntax of the language.

The study also demonstrated that the model could follow complex instructions. Scientists could ask the program to generate a stable RNA sequence that included a specific motif for cloning or excluded a particular pattern that might cause problems. The model successfully adhered to these constraints while still optimizing for stability. This ability to follow natural language commands and produce functional biological parts suggests a new way of engineering life. Instead of relying on slow, expensive cycles of trial and error in the lab, scientists can now use a unified system to predict how RNA will behave and design new molecules to meet precise needs. The work establishes that a single, large-scale model can bridge the gap between understanding biological data and creating new, functional biological tools, opening the door to more efficient design of therapies and research tools.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →