← Latest papers
🧬 biology

The Complex Geometry of Biological Knowledge in Artificial Intelligence: Manifolds, Spectra, and Concept Structure in Foundation-Model Representations

This study reveals that while protein language models encode biologically meaningful information that is linearly decodable, they lack the complex low-dimensional manifolds, multi-scale spectral structures, and hierarchical concept geometries found in single-cell foundation models, instead representing knowledge through a geometrically simple, high-linear-dimensional space shaped by the discrete and homology-driven nature of sequence data.

Original authors: Liu Chen

Published 2026-07-23
📖 8 min read🧠 Deep dive

Original authors: Liu Chen

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Imagine you are trying to understand how a giant library organizes its books. In the world of artificial intelligence, there are "foundation models"—massive computer programs trained on oceans of data to learn the rules of a specific subject. One popular type of these models is trained on biological data, like the genetic code of cells or the sequences of proteins. Scientists have recently discovered that some of these models, specifically those trained on cell data, seem to have a very fancy internal map. They appear to organize information along smooth, curved paths (called "manifolds") and have a complex, multi-layered structure that looks like a beautiful, intricate sculpture. This is exciting because if a computer organizes knowledge like a complex sculpture, it might be easier for us to understand how it thinks and even use that structure to discover new biology.

But here is the big question: Do all biological AI models organize their knowledge this way? Specifically, what about the models trained on proteins? Proteins are the tiny machines that do most of the work inside our bodies, and their "language" is a long string of amino acids. If these protein models also have that fancy, curved, complex geometry, it would be a huge win for scientists trying to decode them. If they don't, it means we might be looking for a treasure map where there is only a flat, open field. This paper sets out to take a magnifying glass to these protein models and see if they really hold the same secret, complex architecture as their cell-focused cousins, or if their internal world is built on a different kind of logic entirely.


The Great Protein Map Hunt: A Tale of Curves vs. Clouds

Meet the Protein Language Models (pLMs). Think of them as super-smart chefs who have tasted every recipe in the universe (hundreds of millions of protein sequences) and learned the rules of cooking without ever seeing a single finished dish. They are so good that they can predict how a protein will fold into a 3D shape or what it does in the body. Scientists have been asking: "How does this chef organize their knowledge in their brain?"

For a while, the answer seemed to be "in a complex, curved maze." Other biological AI models (trained on single cells) were found to have knowledge stored in smooth, low-dimensional curves, like a ribbon winding through a high-dimensional room. They had "spectral" structures (like distinct bands of color in a rainbow) and "concept hierarchies" (like a family tree where ideas branch out neatly). This led to a big hypothesis: Maybe all biological AI organizes knowledge in these fancy, complex geometric shapes.

This paper, led by researcher Liu Chen, decided to put that hypothesis to the test. They built a "three-lens battery" to inspect the brains of eight different protein models, including the famous ESM-2 family (ranging from tiny 8-million-parameter models to massive 3-billion-parameter giants) and others like ProtBERT and Ankh. They didn't just look; they used a set of rigorous tools to measure three specific things:

  1. The Shape (Manifold): Is the data organized on a smooth, curved surface?
  2. The Spectrum: Does the data have distinct, separated bands of information like a rainbow?
  3. The Concepts: Are ideas arranged in a neat hierarchy or family tree?

They also ran these same tests on "synthetic controls"—fake data they created that was known to have complex shapes (like a twisted donut or a spiral). If their tools worked, they should find the complex shapes in the fake data.

The Verdict: It's a Flat, Linear Cloud, Not a Curved Maze

The results were a bit of a surprise, but they were very clear. The authors found that protein language models do not have the complex, curved geometry that cell models have. Instead, their internal knowledge is "real but simple."

1. The Shape: A Cloud, Not a Ribbon
When the researchers tried to find a smooth, curved surface (a manifold) where the protein data lived, they came up empty.

  • The Finding: While the data looks like it might be on a low-dimensional curve (the "nonlinear intrinsic dimension" was small, around 26 for the big models), the reality is different. The data actually spreads out across a huge, high-dimensional space (linear dimension around 131).
  • The Analogy: Imagine trying to flatten a crumpled piece of paper. If it were a smooth ribbon, you could flatten it easily. But these protein models are more like a giant, fluffy cloud of cotton candy. It looks small from a distance, but if you try to squeeze it into a flat shape, it just spreads out.
  • The Proof: When the researchers tried to use fancy, curved math to visualize the data (like UMAP, which is popular for making pretty 2D pictures), the results were terrible. The models lost their ability to predict protein properties. However, simple, straight-line math (linear probes) worked perfectly. The authors conclude that the "geometry" of protein knowledge is not a smooth, curved manifold, but a "near-linear, high-dimensional cloud."

2. The Spectrum: One Smooth Slide, Not a Rainbow
Next, they looked at the "spectrum" of the data, which is like looking at the frequencies of sound in a song.

  • The Finding: They found a single, smooth slide of data following a power law (a specific mathematical curve where the exponent is close to 1). There were no distinct "bands" or gaps that separated different biological concepts.
  • The Analogy: If the cell models were like a rainbow with distinct red, orange, and blue bands, the protein models are like a single, smooth gradient of gray. There are no sharp lines separating one type of protein from another in the data's internal structure.
  • The Nuance: While there was no complex structure, the steepness of that smooth slide (the exponent) was actually very useful. It acted like a "quality score." The flatter the slide (the closer the exponent was to 1), the better the model performed on real-world tasks. So, while the structure wasn't complex, it was a reliable indicator of how good the model was.

3. The Concepts: Flat, Not Hierarchical
Finally, they checked how the models organized "concepts" like "is this a membrane protein?" or "what is its secondary structure?"

  • The Finding: The models were great at answering these questions using simple, straight lines. If you asked a linear probe, "Is this a membrane protein?" it said "Yes" with high accuracy. However, the relationships between these concepts were flat.
  • The Analogy: Imagine a library. In a complex hierarchy, books are arranged by continent, then country, then city, then street. In these protein models, the books are just scattered on a giant, flat table. You can find the right book easily (linear decodability), but there is no neat family tree connecting them. The "analogy" test (checking if the model understands that "A is to B as C is to D") also failed; the relationships were weak and didn't line up in parallel directions like they do in language models.
  • The Confound: The authors also found that some of the apparent structure was actually just the model reacting to basic features like the length of the protein or how many of certain amino acids it had, rather than deep biological meaning.

Why the Difference? The Nature of the Data

The paper offers a fascinating reason for why protein models are different from cell models.

  • Cell Data: Cells change continuously. A stem cell slowly turns into a blood cell. This creates a smooth, winding path (a manifold) in the data.
  • Protein Data: Proteins are discrete. They are made of a string of 20 possible amino acids. The space of all possible proteins is huge and "chunky," organized by evolutionary history (homology) rather than smooth transitions.
  • The Conclusion: Because the data itself is discrete and scattered in clusters, the AI model doesn't need to build a smooth, curved map. It just spreads its knowledge out in a high-dimensional, linear cloud. The "complex geometry" reported in cell models doesn't transfer to proteins because the underlying reality of proteins is different.

The Takeaway

The authors conclude that biological knowledge in protein models is real and accessible, but geometrically simple.

  • What works: Simple, linear tools (like straight-line probes and PCA) are the best way to read these models.
  • What doesn't work: Trying to force these models into complex, curved shapes or looking for deep, hierarchical family trees in their internal structure is a dead end.
  • The Silver Lining: The fact that the knowledge is linear and simple is actually a good thing. It means we can trust these models to give us straight answers without needing to decode a mysterious, curved maze. The "complex geometry" hypothesis, while beautiful for cell models, simply doesn't fit the protein world. The paper confirms that for proteins, the map is not a winding ribbon, but a vast, open, and surprisingly straightforward field.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →