Where in the spectrum does biology live? A confound audit and a compaction law for gene embeddings of single-cell foundation models
This paper reveals that the leading spectral axes of single-cell foundation models are predominantly confounded by gene expression abundance rather than biological meaning, demonstrating that removing these axes improves retrieval performance and establishing a model-dependent compaction law that guides optimal dimensionality reduction for biological tasks.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
The Library of Life and the Loudness of Noise
Imagine you are trying to teach a super-smart robot to understand the secret language of life. To do this, scientists have built "foundation models"—massive AI brains trained on millions of tiny snapshots of cells. These snapshots, called single-cell RNA data, show which genes are active in a cell at a specific moment. Just like a language model learns that the word "the" appears more often than "zebra," these biology models learn that some genes are "loud" (expressed in huge amounts) and others are "whispers" (rarely seen).
The big question researchers have been asking is: When these models learn to represent genes as lists of numbers (called embeddings), are they actually learning the deep, complex rules of biology? Or are they just learning a shortcut? In the world of language, we know that the most obvious patterns in AI often just reflect how frequently a word appears, not what it means. If a biology AI is doing the same thing—just memorizing which genes are loud rather than understanding how they work together—then we might be fooling ourselves into thinking we've discovered deep biological secrets when we've only found a frequency chart. This paper steps in to audit those secrets, asking if the AI is truly smart or just a very good accountant of volume.
The Great Gene Audit: Is the AI Listening or Just Counting?
In this study, researcher Liu Chen acts like a forensic accountant for eight different "single-cell foundation models." These models are like eight different students who have all read the same massive library of cell data and are now trying to write a report on how genes relate to one another. The author wants to know: Are these students actually understanding the story, or are they just highlighting the loudest voices in the room?
The "Loudness" Problem
The paper starts with a simple observation. In these models, a gene's "voice" is determined by two things: how much of it is produced (abundance) and how many cells have it (detection breadth). The author suspects that the most powerful "directions" in the AI's brain—the top lines of its internal map—are mostly just recording this loudness.
To test this, the author compares the AI's gene maps against a "negative control": a model called ESM2. This model is special because it was trained only on protein sequences (the building blocks of genes) and has never seen a single number representing how much a gene is expressed. It's like a student who studied the dictionary but never heard anyone speak.
The Findings: The Top of the Map is Just Noise
The audit reveals a startling pattern. For the models trained on cell data, the very first few "directions" in their brain are overwhelmingly dominated by gene loudness.
- The Evidence: For six out of seven models trained on cell counts, the top axis is 51% to 76% explained by how abundant a gene is. It's as if the AI's first thought is, "This gene is loud," rather than "This gene builds a ribosome."
- The Contrast: The protein-only model (ESM2), which never saw expression numbers, has almost zero connection to loudness in its top axis (only 0.9%). This proves that the "loudness" signal isn't a natural property of genes; it's a side effect of how the other models were trained.
The "All-But-The-Top" Surprise
Here is where it gets counterintuitive. In many AI systems, the most important information is at the very top. But in these biology models, the top is actually a distraction.
- The Experiment: The author tried deleting the top few directions (the "loud" ones) and asked the AI to find relationships between genes again.
- The Result: Surprisingly, the AI got better at finding real biological connections (like which genes work together in a protein complex) after deleting the top directions. The "loudness" was actually getting in the way of the real signal.
- Where Biology Lives: The real biological secrets aren't at the very top or the very bottom. They live in the middle, specifically in the spectral band between axes 32 and 64. It's like finding the best conversation in a noisy party: you have to tune out the loudest shouters (the top axes) and the quietest whispers (the bottom axes) to hear the actual group discussion in the middle.
The "Compaction" Law: One Size Does Not Fit All
The paper also investigates a popular idea that you can shrink these massive AI models down to a small size (like keeping only the top 64 numbers) without losing much information. A previous study suggested that rank 64 was a magic number where you could compress the data and still keep the biology.
The author tests this across all eight models and finds that there is no single magic number.
- The Reality: The number of directions needed depends entirely on the specific model and the specific relationship you are looking for. For some relationships, like genes that are just "co-expressed" (happening to be loud at the same time), you only need a tiny slice (around 24–40 directions). But for complex relationships like "regulator to target" (where one gene controls another), you might need up to 384 directions in some models.
- The Verdict: The idea that "rank 64" is a universal rule is false. While rank 64 is a decent "safe default" that keeps about 85–95% of the information for most tasks, it is not perfect. If you are trying to study complex protein groups or gene regulation, you might be throwing away crucial details by stopping at 64.
What This Means for the Future
The paper concludes with a set of practical, low-cost rules for anyone using these models:
- Check the Volume: If someone claims a specific axis in an AI model represents a biological truth, they must first prove that the axis isn't just measuring how loud the gene is.
- Tune Your Cutoff: Don't just pick a random number like 64 to shrink your model. You need to test how many directions your specific task actually needs.
- The Middle is Gold: If you are looking for deep biological structure, don't look at the very top of the AI's brain. Look in the middle, around axes 32 to 64, where the real signal hides away from the noise of abundance.
In short, these powerful biology AIs are incredibly smart, but they are also easily distracted by volume. To hear the true story of life, we have to learn to ignore the shouting and listen to the middle ground.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.