← Latest papers
📄 molecular biology

FlexRibbon: Joint Sequence and Structure Pretraining for Protein Modeling

FlexRibbon is a novel protein foundation model that jointly pretrains on amino acid sequences and 3D structures using a combination of masked language modeling and diffusion-based denoising, achieving state-of-the-art performance across diverse tasks—particularly in mutation-rich regimes where traditional MSA-based methods struggle—without relying on multiple sequence alignments.

Original authors: Zhu, J., Shi, Y., Bi, R., Jin, P., Liu, C., Zhang, Z., Huang, H., Guo, Z., Hu, P., Ju, F., Huang, L., Tai, X., Li, C., Gao, K., Wei, X., Xia, H., Zhang, J., Min, Y., Wang, Z., Wang, Y., He, L., Liu, H
Published 2026-07-09
📖 5 min read🧠 Deep dive

Original authors: Zhu, J., Shi, Y., Bi, R., Jin, P., Liu, C., Zhang, Z., Huang, H., Guo, Z., Hu, P., Ju, F., Huang, L., Tai, X., Li, C., Gao, K., Wei, X., Xia, H., Zhang, J., Min, Y., Wang, Z., Wang, Y., He, L., Liu, H., Qin, T.

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Imagine the world of proteins as a massive, bustling city. For a long time, scientists trying to understand this city had to choose between two very different maps.

The first map was like a textbook of grammar. It only looked at the sequence of letters (amino acids) that make up a protein, ignoring what the city actually looks like. It was great at understanding the "language" of life but often got lost when trying to figure out the 3D shape of the buildings.

The second map was like a crowdsourced survey. It relied on looking at thousands of similar cities (evolutionary cousins) to guess what a new one would look like. This was incredibly accurate for standard buildings, but if you tried to map a brand-new, weirdly shaped skyscraper or a city that had changed its layout too much, the survey failed because there were no similar examples to compare it to.

Enter FlexRibbon, a new kind of mapmaker that decided to stop choosing between the textbook and the survey. Instead, it learned to read the grammar and look at the blueprints at the same time.

The Secret Sauce: A Two-Way Street

Most previous models were like a one-way street: they took a sequence of letters and tried to guess the shape. FlexRibbon built a two-way street. It learned that the letters determine the shape, but the shape also tells you what the letters should be.

To teach this, the researchers used a clever training game. They took a protein, scrambled its 3D shape with noise (like shaking a snow globe), and asked the model to clean it up. At the same time, they hid some of the letters in the sequence and asked the model to guess them, but with a twist: it had to use the messy 3D shape as a hint to make the right guess. This forced the model to understand that the sequence and the structure are locked together, like a lock and key that can't be separated.

Why This Matters: The "Messy" Parts

The real magic happens in the messy parts of the city. In biology, some proteins have flexible, floppy arms (like antibody loops) or change shape when they grab onto something. The old "survey" models (which rely on finding similar cousins) often get stuck here because these floppy parts change too much to have a clear pattern in the evolutionary data.

FlexRibbon, however, handles these tricky regions differently. Instead of relying on a "crowd of cousins" to find a pattern, it bypasses the need for that survey entirely. By directly modeling the link between the individual sequence and its structure, it can predict these flexible, highly mutated regions with surprising accuracy, even when evolutionary signals are weak or missing.

The Proof: 12 Different Challenges

The team didn't just say it worked; they put FlexRibbon through a gauntlet of 12 different tasks to see if it could handle real-world problems. Here is what they found:

  • Antibody Design: When asked to design antibodies (the body's defenders) to stick to viruses, FlexRibbon succeeded 61.3% of the time for standard antibodies and 51.1% for nanobodies (tiny antibodies). This was a big jump over the previous best methods, which were stuck around 46.7% and 44.0% respectively.
  • Peptide Binding: For predicting how small protein chains (peptides) stick to larger proteins, FlexRibbon hit a success rate of 91.4%, beating the next best method by 7.0% to 10.2%.
  • Drug Discovery: When trying to guess how well a drug molecule would stick to a protein, FlexRibbon achieved a correlation score of 0.848 and an error rate of 1.150, outperforming all other methods tested.
  • Shape-Shifting: The model could also predict how a protein changes shape when a drug binds to it, achieving a structural similarity score (TM-score) of 0.985 for one test case, which is incredibly close to the real thing.

What It's NOT

It's important to know what FlexRibbon doesn't do. It doesn't rely on the "crowdsourced survey" (Multiple Sequence Alignments) that older models used to make predictions. If you give it a protein with very few evolutionary cousins, the old models might struggle because their survey has no data, but FlexRibbon succeeds by relying on its internal understanding of the direct sequence-structure link, allowing it to navigate regions where alignment signals are weak.

The Bottom Line

The paper suggests that by teaching a model to learn sequences and structures together, rather than one after the other, we get a much more flexible and accurate tool. It's not just a better guesser; it's a better designer. Whether it's creating new antibodies to fight diseases or figuring out how drugs fit into their targets, FlexRibbon shows that when you let the sequence and the structure talk to each other, the whole system works better.

The authors measured these results on specific benchmarks like the CASP15 challenge and the PoseBusters V1 dataset, finding that FlexRibbon consistently hit new high scores, especially in the tricky, mutation-heavy areas where other methods often stumble. It's a step toward a future where we can design proteins from scratch with the same ease as writing a sentence.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →