In Search of Manifolds Inside Protein Language Models: A Largely Negative Report
This paper systematically investigates the manifold hypothesis in protein language models (pLMs) using a comprehensive battery of geometric and topological diagnostics, concluding that unlike single-cell foundation models, pLM representations are predominantly high-dimensional and linear rather than organized around low-dimensional manifolds.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
Imagine you are trying to understand a massive, chaotic library. In this library, every single book is a protein, the tiny molecular machines that build and run all living things. For a long time, scientists have been using super-smart computer programs, called "protein language models," to read these books. These programs are like digital detectives that have read billions of protein sequences and learned to predict how they fold, what they do, and how they might change.
But here is the big question: How does the computer actually "think" about these proteins? When the computer looks at a protein, does it see a messy pile of data, or does it see a hidden, smooth shape? Scientists have a favorite idea called the "manifold hypothesis." Think of it like this: imagine a giant, crumpled piece of paper floating in a huge 3D room. Even though the paper is floating in 3D space, if you were an ant walking on it, the world would feel flat and two-dimensional. The "manifold hypothesis" suggests that even though the computer's brain is huge and complex (high-dimensional), the important information about proteins might actually be hiding on a smooth, low-dimensional surface, like that flat piece of paper. This idea has been a hit in other fields, like studying how cells grow and change, where scientists found these smooth, low-dimensional "roads" that cells travel along. So, the big question is: Do these protein-detecting computers also have these smooth, hidden roads inside their brains?
A researcher named Olivia Denvis decided to go on a treasure hunt to find these hidden roads inside protein language models. She didn't just guess; she built a massive "detection battery," which is like a super-advanced metal detector and map-maker rolled into one. She tested this tool on several different protein models, including the famous ESM-2 family (which comes in sizes from tiny 8 million parameters to a giant 3 billion) and others like ProtBERT and ProtT5. She used a variety of clever tricks to see if the data was actually sitting on a smooth, low-dimensional surface. She checked if the data had a specific "intrinsic dimension" (a fancy way of asking: how many directions do we really need to describe this?), she tried to flatten the data onto 2D maps to see if patterns emerged, and she even looked for topological loops (like the hole in a donut) that would prove the data was shaped like a specific object.
The result of this hunt? A largely negative report. The treasure hunt came up empty.
Denvis found that the idea of a smooth, low-dimensional road inside these protein models is probably wrong. Instead of finding a flat sheet of paper floating in a 3D room, she found that the data is actually spread out across a vast, high-dimensional space. When she measured the "intrinsic dimension" (the number of directions needed to describe the data), it wasn't a small, manageable number like 2 or 3. For the medium-sized model (ESM-2 650M), the dimension was around 63, and for the giant 3-billion-parameter model, it jumped to 88. Even more telling, as she made the models bigger and smarter, the dimension got bigger, not smaller. If there were a single, smooth road, the dimension should have stayed the same or gotten simpler as the models improved. Instead, the models seemed to be using more and more directions to store information, suggesting the structure is complex and "high-dimensional" rather than a simple curve.
She also tried to use the computer's "map-making" tools to flatten the data down to 2D, just like we do with single-cell data. But when she did this, the useful information got destroyed. It was like trying to flatten a complex, multi-layered cake into a pancake; the flavor (the biological information) got lost. When she tested these flattened maps to see if they could predict things like protein stability or structure, they performed terribly. In fact, the simplest method—just looking at the raw, high-dimensional data with a straight line (a linear probe)—worked the best. The computer didn't need a curved, low-dimensional road; it needed a wide, flat highway with dozens of lanes.
Furthermore, she looked for "loops" or holes in the data, which would indicate a specific shape like a donut or a circle. She found none. The data was topologically "trivial," meaning it was just a big, solid blob without any interesting shapes. She also discovered that the clusters of proteins that looked like they were forming neat groups on 2D maps were actually just an illusion. When she removed the influence of simple things like how long the protein was or what basic ingredients (amino acids) it was made of, the groups fell apart. The "manifold" people thought they saw was just a reflection of these basic, boring facts, not a deep, hidden structure.
To make sure her tools weren't broken, she tested them on things she knew had these smooth roads, like a synthetic "Swiss roll" shape and real data from single cells. Her tools worked perfectly there, finding the low dimensions and the loops. This proved that her methods were good; the problem wasn't the tools, it was the protein data itself.
So, what does this mean? It turns out that protein language models are not like the smooth, continuous maps of cell development. Instead, they are more like a vast, high-dimensional library where every book is organized by a complex mix of length, composition, and family history. The information is there, and it's incredibly useful, but it's not hiding on a simple, low-dimensional surface. It's spread out across a high-dimensional space, and trying to force it into a simple 2D map actually throws away the most important details. The paper concludes that for protein models, we should stop looking for smooth, curved roads and start appreciating the wide, high-dimensional highways they actually travel on.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.