← Latest papers
🤖 machine learning

Proteo-R1: Reasoning Foundation Models for De Novo Protein Design

Proteo-R1 is a dual-expert framework that enhances *de novo* protein design by explicitly decoupling molecular reasoning from geometric generation, using a multimodal language model to identify critical functional residues as hard constraints for a diffusion-based generator to ensure stable, interpretable, and controllable design outcomes.

Original authors: Fang Wu, Weihao Xuan, Heli Qi, Hanqun Cao, Heng-Jui Chang, Zeqi Zhou, Haokai Zhao, Ma Jian, Carl Ma, Yu-Chi Cheng, Kuan Pang, Xiangru Tang, Zehong Wang, Guanlue Li, Hanchen Wang, Kejun Ying, Pan Lu, C
Published 2026-05-06
📖 4 min read☕ Coffee break read

Original authors: Fang Wu, Weihao Xuan, Heli Qi, Hanqun Cao, Heng-Jui Chang, Zeqi Zhou, Haokai Zhao, Ma Jian, Carl Ma, Yu-Chi Cheng, Kuan Pang, Xiangru Tang, Zehong Wang, Guanlue Li, Hanchen Wang, Kejun Ying, Pan Lu, Chiho Im, Seungju Han, Peng Xia, Tinson Xu, Yinxi Li, Deyao Zhu, Pheng-Ann Heng, Naoto Yokoya, Masashi Sugiyama, Li Erran Li, Jure Leskovec, Yejin Choi

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a master architect trying to design a new, custom-shaped key that fits perfectly into a specific, complex lock (a protein target).

The Old Way: The "Guess-and-Check" Architect
Previously, AI models trying to design these keys worked like a frantic artist throwing paint at a canvas. They would generate the entire shape of the key all at once, hoping the atoms landed in the right places. While they were getting better at making the key look real, they weren't really "thinking" about why certain parts of the key needed to be a specific shape. They were just sampling random possibilities until something looked okay. This made it hard to understand why the AI made a specific design choice, and if the key didn't work, it was hard to fix just one part without ruining the whole thing.

The New Way: Proteo-R1 (The "Strategist" and the "Builder")
The paper introduces Proteo-R1, which changes the game by splitting the job into two distinct roles, much like a construction project with a Strategist and a Builder.

  1. The Strategist (The Reasoning Expert):

    • Who they are: A super-smart AI that reads the "blueprint" (the protein sequence and structure) and the "client's notes" (what the key needs to do).
    • What they do: Before drawing a single line, this AI stops to think. It identifies the most critical spots on the lock—the "hotspots" where the key must touch to work. It decides, "Okay, at this specific point, we need a charged anchor here, and a hydrophobic spot there."
    • The Analogy: Think of this as the architect saying, "We need a steel beam right here to hold the roof, and a window right there for light." They make the hard, logical decisions first.
  2. The Builder (The Generation Expert):

    • Who they are: A highly skilled 3D printer (a diffusion model) that is amazing at creating smooth, realistic shapes.
    • What they do: The Strategist hands the Builder a set of strict rules: "Build the rest of the key, but you must keep this steel beam here and that window there." The Builder then fills in the rest of the design, ensuring the atoms fit together perfectly while respecting the Strategist's non-negotiable rules.
    • The Analogy: The Builder is like a master mason who knows exactly how to lay bricks to make a beautiful wall, but they are told exactly where the door and windows must go. They don't guess; they just execute the plan perfectly within those constraints.

Why This is a Big Deal

  • No More "Black Box": In the old way, you couldn't tell why the AI picked a certain shape. With Proteo-R1, you can see the Strategist's notes: "I put this amino acid here because it forms a salt bridge." It's transparent and explainable.
  • Better Control: If you want to change the design, you don't have to throw away the whole thing. You just tell the Strategist to change one decision (e.g., "Move the window"), and the Builder adjusts the rest.
  • Realism: Because the Builder is a top-tier geometric model, the final keys look physically real and don't have atoms crashing into each other (which happens when you just guess).

How They Trained It
The paper describes a three-step training camp for this team:

  1. Alignment: Teaching the Strategist to speak the same language as the Builder, so they understand protein structures and sequences correctly.
  2. Reasoning Drills: Giving the Strategist thousands of puzzles to solve, like "Find the salt bridges" or "Locate the hotspots," so it gets really good at identifying critical parts of a protein.
  3. Joint Practice: Finally, they practice together. The Strategist makes a plan, the Builder tries to build it, and if the building fails, the Strategist learns from that mistake to make a better plan next time.

The Results
When they tested this on designing antibodies (a type of protein key used by the immune system), Proteo-R1 created keys that fit their locks better and more realistically than previous methods. Crucially, it didn't just copy old keys; it designed new ones that were structurally sound because it understood the logic of the design, not just the shape.

In short, Proteo-R1 stops AI from blindly guessing and starts it by reasoning first, then building, just like a human expert would.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →