Single-cell foundation models benefit from cross-modal training: adding proteomics data beats parameter scaling
This paper demonstrates that cross-modal pretraining of a single-cell foundation model with proteomics data yields superior gene-level and cell-level representations and better generalization than simply scaling up RNA-only models, highlighting the value of multimodal training over parameter scaling alone.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
Inside every living cell, a complex conversation is taking place, written in two different languages. One language is made of RNA, molecules that carry instructions for building proteins, the workhorses that keep a cell alive and functioning. Scientists have spent years learning to read these RNA messages, creating massive digital libraries that describe how cells behave in health and disease. These libraries have grown so large that researchers now use powerful computer programs, known as foundation models, to find patterns within them. These programs learn to recognize the unique signatures of different cell types, much like a librarian who can instantly identify a book's genre by its cover. For a long time, the prevailing belief was that the only way to make these programs smarter was to feed them more and more RNA data, hoping that sheer volume would eventually reveal every secret of cellular life.
However, cells speak more than one language. Alongside RNA, they produce proteins, the actual structures that perform the cell's tasks. While RNA tells the cell what to build, proteins are the finished products. Until recently, most computer models of cells ignored these proteins, focusing exclusively on the RNA instructions. This left a gap in understanding, as the instructions do not always match the final outcome. The question facing scientists was whether adding this second layer of information—the protein data—could teach these computer models more than simply adding more RNA data ever could. It was a test of whether quality and variety in information could outperform the simple strategy of making the models bigger and feeding them more of the same.
A team of researchers set out to answer this by training a computer model on a new kind of data. They started with an existing model that had already learned from hundreds of millions of RNA samples. Instead of just feeding it more RNA, they introduced it to a vast collection of protein profiles. These profiles came from 440 different studies that used mass spectrometry, a technique that measures the specific proteins present in a cell. The researchers carefully guided the model to learn from these 48,843 protein samples, a process they called cross-modal continued pretraining. This meant the model had to learn how the protein language related to the RNA language it already knew, effectively teaching it to understand the cell through both its instructions and its finished products.
The results were striking. The researchers trained a model with 70 million parameters, a measure of its complexity, on this mixed data for just one pass through the protein samples. When they tested this model, it performed as well as, or better than, models that were vastly larger and trained only on RNA. Specifically, the smaller model with protein data matched or exceeded the performance of RNA-only models that had 1 billion and 3 billion parameters. This finding challenges the idea that the only path to better understanding is to simply scale up the size of the model and the amount of RNA data. Instead, it suggests that introducing a different type of biological data can provide a much more efficient boost to the model's intelligence.
The benefits of this approach extended beyond general performance. The model trained with protein data showed a stronger ability to generalize, meaning it could apply what it learned to new situations it had never seen before. This was particularly clear when the model was tested on how cells respond to protein changes. In these tests, simply making the RNA-only model bigger did not help it understand these changes any better. The model that had learned from proteins, however, handled these new challenges with greater ease. This indicates that the protein data helped the model build a more robust and flexible understanding of how cells work, one that is not tied strictly to the patterns found in RNA alone.
These findings suggest that the future of understanding cells may not lie in building ever-larger models trained on a single type of data. Instead, the most powerful insights may come from carefully combining different kinds of biological information. By teaching computer models to read both the instructions and the products of the cell, scientists can create tools that are not only smarter but also more accurate in their predictions. This approach offers a promising path toward more informative models that can capture the full complexity of life, proving that sometimes, adding a new perspective is far more valuable than simply adding more of the same.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.