Geometric Representations of Knowledge Inside Biological Large Language Models: an Empirical Analysis
This empirical study reveals that while biological large language models (SCFMs) do encode a real, low-dimensional, and partially linearized geometric structure for biological knowledge, this structure is modest, often nonlinearly entangled, and largely overlaps with what classical expression-based methods like PCA can already recover, falling short of the crisp, regular geometry observed in natural language models.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
Imagine you have a giant, magical library where every book is a single living cell from the human body. For years, scientists have tried to read these books to understand how our bodies work, but the text is messy and hard to decipher. Recently, a new kind of "super-reader" called a Biological Large Language Model (or Biological LLM) was invented. These are computer programs trained on millions of cell "books" to spot patterns, much like how your phone's keyboard learns to predict the next word you'll type.
The big question scientists have been asking is: Do these super-readers organize knowledge in a neat, straight-line way? In regular language models (the ones that write essays or chat with you), researchers found that ideas like "king" and "queen" sit on a straight line, and opposites are just a simple step away from each other. It's like a perfectly organized map where every concept has its own clear street. Scientists hoped these biological super-readers would have the same neat, straight-line map for things like "healthy vs. sick" or "growing vs. resting." If they did, it would mean we could easily tweak the computer's thinking to understand diseases or development. But, as we are about to see, the map inside these biological models is a bit more like a winding, hilly forest path than a straight city street.
The Map Inside the Machine: A Forest, Not a Highway
In this study, a researcher named Olivia Denvis decided to take a deep dive into four of these biological super-readers (named scBERT, Geneformer, scGPT, and UCE) to see if they really have that neat, straight-line organization. She treated them like explorers, sending them into a massive dataset of 420,000 cells with detailed notes on what they are, where they live, and what they are doing. She then compared the models' internal "maps" against the old-school, reliable tools scientists have used for decades, like simple math tricks called PCA.
Here is what she found: The models do have a map, but it's not the crisp, straight highway everyone hoped for.
1. The Map is Compact, but Crowded
First, the researchers checked how much space the models use to store information. Imagine a huge warehouse (the model's full size) that could hold 1,280 shelves. The study found that the models only actually use about 12 to 26 shelves to store their knowledge. That's a very low-dimensional space, which is good because it means the models are efficient. However, the space is "anisotropic," which is a fancy way of saying it's shaped like a long, narrow cone rather than a round ball. The smaller models were even more squeezed into this narrow cone, suggesting they might be a bit "cramped" in how they think.
2. The Easy Stuff is Straight, The Hard Stuff is Curvy
When the researcher asked the models to identify simple things, like "Is this cell from a male or female?" or "Is it in the immune system?", the models were amazing. They could draw a straight line to separate these groups, just like the language models do. In fact, for these easy categories, the models were almost perfect.
But here is the twist: When she asked about the interesting stuff—like "Is this cell sick?" or "What specific subtype is this?" or "How far along is it in its development?"—the straight lines stopped working. The knowledge was still there, but it was tangled up in curves.
- The Proof: When the researcher tried to use a simple straight line to predict a cell's development time, it only got about 61% of the answer right. But when she let the computer follow the natural, curved path of the data (like walking along a winding trail instead of cutting through a field), the accuracy jumped to 80%.
- The Takeaway: The models know the complex stuff, but they store it in a curved, non-linear way that simple straight lines can't easily reach. This is different from language models, where even complex ideas often sit on straight lines.
3. The "Steering Wheel" is a Bit Wobbly
In language models, if you want to change a word's meaning (like turning "happy" into "sad"), you can just add a specific vector (a direction) to it, and it works perfectly. The researcher tried this with the biological models. She tried to "steer" a cell's identity by adding a direction vector.
- The Result: It worked, but only about 54% to 66% of the time. It was better than a coin flip, but far from the perfect control seen in language models. Sometimes, trying to change one thing accidentally changed something else. The directions for different ideas were somewhat parallel, but not perfectly so, and they were only weakly aligned.
4. The Models Agree with Each Other, Not the "Truth"
The researcher also checked if all four models were thinking the same way. They were moderately similar to each other (about 55% to 74% similar), which is nice. But when she compared them to the "Gold Standard" of biological knowledge (a detailed map of how cell types are related in real life), the models were only about 36% to 45% similar.
- The Surprise: The models actually looked more like the old-school math tools (PCA) than they looked like the real biological truth. They converged on a shared, simplified view of the data that was closer to the basic math than to the complex reality of biology.
5. The "Deep" Knowledge is in the Middle
Finally, she looked at where in the model's layers (its "brain") this knowledge lived.
- Simple Identity: Knowing "what kind of cell" this is gets stronger the deeper you go into the model.
- Disease Status: Knowing if a cell is sick peaked in the middle layers and then faded away as the model got deeper.
- Batch Effects: The models were good at ignoring technical noise (like which machine took the picture) in the early layers, but they didn't completely forget it.
The Bottom Line
So, what's the verdict? Biological Large Language Models are real, and they do hold a geometric map of knowledge. But it's not the clean, straight-line, perfectly organized map we saw in language models. Instead, it's a "partially linearized, low-dimensional" map.
The models are great at the easy stuff (like telling a male cell from a female one), but for the hard, clinically important stuff (like disease or development), the knowledge is curved and tangled. Most of the "straight-line" knowledge the models have is actually just what we could already get from simple, old-school math tools. The new stuff the models learned is there, but it's hidden in the curves, making it harder to read and use.
The study concludes that while these models are useful, they aren't the magic "straight-line" breakthrough some hoped for. The real challenge for the future isn't just building bigger models, but building better tools to read the curved, winding paths where the most important biological secrets are actually hiding.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.