← Latest papers
💻 bioinformatics

A geometric atlas of how ESM3 organizes modalities across depth

This study reveals that in the multimodal protein language model ESM3, distinct physical modalities (sequence, structure, secondary structure, and solvent accessibility) initially occupy separate subspaces before fusing into a shared low-dimensional representation between layers 25 and 35, while functional annotations remain orthogonal throughout, with this fusion process being a learned, universal property independent of protein length but delayed by structural disorder.

Original authors: Steenwyk, J. L.

Published 2026-07-12
📖 5 min read🧠 Deep dive

Original authors: Steenwyk, J. L.

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Imagine a massive, 1.4-billion-parameter digital brain called ESM3. Its job is to understand proteins, the tiny molecular machines that build life. But this brain doesn't just read the protein's "recipe" (its amino acid sequence); it also looks at its 3D shape, its secondary structure (like whether a part is a coil or a helix), how much water it touches, and even a list of what it does (functional annotations).

The big question this paper asks is: How does this brain mix all these different types of information together? Does it keep them in separate drawers, or does it blend them into a smooth smoothie?

The "Layer Cake" of Understanding

Think of ESM3 as a 48-story skyscraper. Information enters at the bottom (Layer 0) and travels up to the top. The authors took a "geometric atlas" (a map of the brain's internal geometry) to see where the different information streams live as they travel up the building.

1. The Great Separation (Layers 0–24)
When the information first enters the building, the different types of data are like guests at a party who refuse to talk to each other.

  • The physical guests (sequence, 3D structure, secondary structure, and water exposure) start in their own distinct corners.
  • The functional guest (the list of what the protein does) is even more isolated. It stays in a completely different room, never mingling with the others.
  • As they travel up the first half of the building, these groups actually get more separated, reaching a peak of isolation around Layer 24. It's as if the building is organizing the guests into strict, non-interacting cliques.

2. The Great Fusion (Layers 25–35)
Then, something magical happens. Between Layer 25 and Layer 35, the physical guests stop ignoring each other and merge into a single, shared dance floor.

  • The Order of Arrival: The three structure-based guests (3D shape, secondary structure, and water exposure) are already best friends from the very start. They hold hands immediately.
  • The Latecomer: The sequence (the amino acid recipe) is the shy one. It stays in the corner until Layer 28, only joining the dance floor after the structure guests have already aligned.
  • The Result: By Layer 35, all four physical modalities have fused into one low-dimensional "cloud" of understanding. They are no longer separate; they are a unified team.

3. The Lone Wolf: Functional Annotations
Here is the most surprising part: The functional annotation (the "what it does" list) never joins the party.

  • Even though the physical modalities fuse into a shared space, the functional data stays in a completely separate, orthogonal (at a 90-degree angle) dimension across all 48 layers.
  • The authors tested if this was because the functional data was "whole-protein" (one label for the whole thing) while the others were "per-residue" (labels for every tiny part). They added per-residue functional data, and guess what? It still stayed separate.
  • This suggests the brain chooses to keep function on a separate track, not because of how the data is formatted, but because of how it learned to organize information.

What This Is (and Isn't)

The authors were very careful to rule out some obvious explanations:

  • It's not just math: If you take the exact same building but give it random, untrained weights (a "randomly initialized model"), the fusion never happens. The guests stay separated forever. This proves the fusion is a learned property, not just a side effect of adding numbers together.
  • It's not an averaging trick: The authors worried that maybe they just "averaged" the data across the protein, which made it look like it fused. They checked individual residues (without averaging), and the fusion still happened. The fusion is real, happening at the level of individual parts, not just the whole.
  • It's not about the protein's length: Whether the protein is short or long doesn't change when the fusion happens. However, disorder does. Proteins with lots of "coil" (disordered parts) fuse a bit later, while well-folded helical proteins fuse earlier.
  • It's universal: They tested this on 5,555 proteins from 12 different organisms, covering bacteria, archaea, and eukaryotes (including humans). Every single organism reached peak fusion at exactly Layer 35. This pattern is locked into the model's design, not just a quirk of human proteins.

The "Shape" of the Fusion

When the fusion happens, the data doesn't just collapse into a single point or spread out randomly.

  • The "effective rank" (a measure of how many dimensions the data uses) drops to about 3 when the groups are most separated, then expands to about 85 dimensions in the fusion zone.
  • The data reorganizes its variance: it stops being different between groups and becomes different within the group.
  • Crucially, the brain never loses the ability to tell which modality is which. Even in the fused zone, a simple probe can still identify the source with 100% accuracy in the full space. The fusion is a controlled reorganization, not a messy blur.

The Bottom Line

The paper suggests that ESM3 has a very specific, ordered way of thinking:

  1. It first builds separate, specialized features for structure and sequence.
  2. It waits until the middle of the network (around Layer 35) to fuse all the physical world data (shape, sequence, etc.) into a unified representation.
  3. It keeps the "functional labels" (what the protein does) on a completely separate, parallel track, never letting them mix with the physical data.

This isn't just a random pattern; it's a learned, universal strategy used by this specific 1.4-billion-parameter model across the entire tree of life. The authors note that while this explains how the model organizes data, they haven't yet proven why this specific arrangement is the best one, or if it holds true for larger, closed-access versions of the model. But for the open version they studied, the map is clear: Structure fuses, sequence joins late, and function stays apart.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →