← Latest papers
🧬 biology

Spectral universality in single-cell foundation models: a low-dimensional shared core carries the biology

Despite significant architectural and parameter differences, single-cell foundation models share a small, robust, and biologically interpretable low-dimensional core that outperforms their individual full embeddings in functional retrieval tasks, revealing that most of their representational capacity is private and biologically inert.

Original authors: Liu Chen

Published 2026-08-06
📖 5 min read🧠 Deep dive

Original authors: Liu Chen

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Imagine you are trying to understand the secret language of life. Inside every cell of your body, there is a massive instruction manual written in a code called RNA. Scientists have built powerful computer programs, known as "foundation models," to read these manuals. Think of these models as super-smart translators that have read millions of cell instructions to learn how genes work together. For a long time, researchers hoped that if they built enough of these translators, they would all eventually agree on the same "truth" about biology, just like different dictionaries eventually agreeing on the definition of a word.

But here is the tricky part: these models are built by different teams, using different math and different training methods. It's like asking a group of chefs who all learned to cook in different countries to describe the "perfect soup." Do they all agree on what makes a soup taste good? Or does each chef have their own secret, unique recipe? This question matters because if the models disagree, then any discovery made by one might be a fluke of that specific computer program, not a real fact about your body. If they do agree, however, we might have found a universal key to understanding how cells function.

This paper sets out to settle the debate by bringing together seven different single-cell foundation models and one protein-language model (a translator that only knows amino acids, not RNA). The researchers asked a simple question: Do these models actually agree with each other?

The answer turned out to be a bit surprising, like finding a hidden treasure map inside a pile of confusing notes. First, the researchers found that the models are mostly strangers. If you compare the internal "thinking space" of two different models, they barely overlap. On a scale where 1 means they are identical and 0.014 means they are just guessing randomly, these models only agreed by about 0.107. In fact, a simple, untrained chart of how genes naturally stick together (called a co-expression profile) was actually more similar to the models than the models were to each other! It's as if the models are all speaking different dialects, and a basic textbook is easier to understand than any single model's unique jargon.

However, the story doesn't end there. The researchers used a special mathematical tool to look for a tiny, shared core hidden deep inside all these different models. They found it! There is a small, low-dimensional "shared core"—just 43 directions of information—that is present in every single model. Even better, this tiny core is where the vast majority of the real biology lives.

Here is the twist: The vast majority of what each model "knows" (its private, unique dimensions) carries almost no retrievable biological information. When tested, these private parts scored at or barely above random chance, meaning they are largely biologically inert. If you take a model and throw away its unique parts, keeping only the shared 43 directions, the model actually gets better at solving biological puzzles. In fact, a tiny 8-direction version of this shared core, which belongs to no single model but is built from the agreement of all of them, beat every single full-sized model in the study. It's like discovering that while each chef has a huge pantry full of random spices (the private parts), they all secretly agree on the exact same pinch of salt and pepper (the core) that makes the soup taste right.

The researchers also checked if this agreement was just because the models were counting how often genes appear (like counting how many times the word "the" appears in a book) or simply copying the training data. They removed those factors, and while the core shrank slightly, a significant biological signal remained. The core is not entirely free of these statistical influences, but it contains genuine learned biology that survives even after removing abundance and co-expression.

So, what does this core actually tell us? It turns out to be a compact map of life's basic programs. The very first shared direction separates genes that are always "on" (like the machinery that keeps the cell breathing) from genes that are almost never used (like the genes for smelling smells, which are turned off in most tissues). The other directions map out major biological themes: cell adhesion (how cells stick together), immune responses, and how cells build proteins.

The big takeaway is that these models are not interchangeable. You can't just take a finding from one model and assume it applies to the others. But, if you look for the tiny slice of information that all of them agree on, you find the most reliable biological truth. The paper suggests that instead of trying to understand one giant, complex model, scientists should focus on the intersection—the small, shared core—where the real biology is hiding. It's a reminder that sometimes, the truth isn't in the noise of one person's opinion, but in the quiet agreement of the whole group.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →