LucaCell: a sequence-centric foundation model for cross-species single-cell analysis
LucaCell is a sequence-centric foundation model that utilizes pre-trained mRNA sequence embeddings instead of fixed gene identifiers to enable robust cross-species, cross-modal, and cross-context single-cell analysis, including accurate cell type transfer, microbial species differentiation, and host-virus interaction prediction.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
Inside every living cell, a complex library of instructions dictates how that cell functions, grows, and responds to its environment. These instructions are written in a chemical language called RNA, which acts as a messenger carrying orders from the cell's master blueprint. For decades, scientists have developed tools to read these messages, allowing them to catalog the different types of cells in the human body and understand how diseases disrupt their normal operations. However, a major hurdle has always remained: these tools rely on fixed names and labels for genes, much like a library that only works if every book has a specific, pre-assigned call number. If a scientist tries to compare a human cell to a mouse cell, or a human cell to a microbe, the old systems often fail because the names do not match, or the data comes from a different type of measurement entirely. This limitation has made it difficult to build a single, universal understanding of life that works across different species and biological contexts.
A team of researchers has now introduced a new approach that bypasses these naming problems entirely. Instead of relying on static labels, they built a computer model that learns directly from the raw sequence of letters that make up the genetic code. By teaching the model to recognize the patterns in the chemical letters themselves, the researchers created a system that can understand the essence of a gene without needing to know its specific name or the species it comes from. This new model, called LucaCell, treats the genetic code as a continuous language rather than a list of disconnected items. The result is a tool that can translate biological information across species, predict how cells will react to changes, and even identify infections before they become obvious, offering a more flexible and powerful way to study the building blocks of life.
The core of this breakthrough lies in how the model represents genes. Traditional methods treat each gene as a unique symbol, similar to a word in a dictionary. If the dictionary changes or if you are reading a book in a different language, the symbols no longer make sense. LucaCell, however, looks at the actual sequence of letters that form the gene. It uses a pre-trained system to convert these sequences into mathematical shapes that capture their meaning and relationships. When the model encounters a gene, it understands its function based on its chemical structure, not its label. This allows it to recognize that a gene in a human kidney and a similar gene in a mouse kidney are related, even if they have different names or if the data was collected using different technologies. The researchers tested this by training the model on a massive dataset of eighty-five million cells from humans and mice, covering dozens of tissues and disease states. This training gave the model a deep, general understanding of how cells behave.
Once trained, the model was put to the test in several challenging scenarios where older systems typically struggle. In one experiment, the researchers asked the model to identify cell types in the kidneys of humans, mice, and a small primate called a gray mouse lemur. The model successfully matched cell types across these different species, even when the data came from different types of measurements, such as reading the genetic code directly versus reading the accessibility of the DNA. It performed particularly well on the unseen mouse data, suggesting that the sequence-based approach captures fundamental biological truths that apply across species. The model also proved capable of working with microbes. In a test involving single bacteria, the model analyzed raw genetic reads without needing to align them to a reference genome first. It successfully distinguished between more than fifty different bacterial species and could even detect subtle changes in the bacteria's internal state, such as their reaction to antibiotic stress, simply by looking at the patterns in their genetic sequences.
The power of this sequence-centric view also extended to predicting how cells respond to specific changes. The researchers used the model to simulate what would happen if a specific gene were turned off or altered, a process known as perturbation. In these tests, the model predicted the resulting changes in gene activity more accurately than previous systems. It could also account for tiny variations in the genetic code between individuals, known as single-nucleotide polymorphisms. By incorporating these small differences, the model improved its ability to predict gene expression levels for specific people, showing that even minor changes in the genetic sequence can have measurable effects on how a cell functions. This level of detail suggests that the model captures nuances that broader, label-based systems often miss.
Perhaps one of the most striking applications was in the realm of host-virus interactions. The researchers trained the model to predict the viral load in individual cells infected with different strains of the influenza virus. When faced with new virus-host combinations that it had never seen before, the model outperformed all other existing tools, correctly identifying infection patterns where others failed. It was even able to look at cells that had not been exposed to the virus and identify a small subpopulation that already displayed the genetic signatures of an impending infection. This finding suggests that the model can detect subtle, pre-existing states in cells that make them more susceptible to infection, a capability that could be crucial for understanding how diseases spread at the cellular level.
The success of LucaCell suggests that the sequence of genetic letters contains a transferable foundation for understanding biology. By moving away from fixed identifiers and focusing on the underlying chemical language, the researchers have created a model that is more adaptable and robust. While the current version relies on human and mouse data and uses a simplified view of the genetic sequence, the results indicate that this approach can bridge gaps between species, data types, and biological contexts. It offers a new way to see the cellular world, one where the fundamental language of life is understood directly, without the need for translation or rigid categorization. This shift from labels to sequences may well become the standard for future discoveries in how cells operate, interact, and respond to the challenges of life.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.