Inverse FoldDir: Structure-conditioned Protein Sequence Design by Dirichlet Flow Matching
The paper introduces Inverse FoldDir, a controllable inverse-folding method based on Dirichlet flow matching that generates diverse, experimentally validated protein sequences with user-defined constraints by performing iterative denoising on the amino acid probability simplex, achieving state-of-the-art structural recovery metrics and successful functional redesign of an anti-GFP nanobody.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
Proteins are the molecular machines that keep life running, folding into intricate three-dimensional shapes to perform tasks like building cells, fighting infections, and digesting food. The secret to their power lies in a simple chain of building blocks called amino acids; the specific order of these blocks determines how the chain twists and turns into a functional structure. For decades, scientists have been able to predict the shape a protein will take if they know its amino acid sequence. But the reverse problem—figuring out which sequence of amino acids will fold into a specific, desired shape—has been a much harder puzzle. This reverse engineering is crucial for designing new medicines, creating better materials, and even building proteins from scratch that nature never evolved. The challenge is that a single protein shape can often be built by many different sequences, and finding the right one requires balancing strict structural rules with the flexibility needed for the protein to actually work in a living system.
A team of researchers has now developed a new method called Inverse FoldDir to solve this puzzle with greater control and accuracy. Instead of guessing the amino acid sequence one piece at a time, which can lead to dead ends, their approach treats the entire protein sequence as a fluid cloud of possibilities that gradually solidifies into a final design. Imagine trying to find the perfect path through a dense forest; rather than picking a single direction and walking until you hit a wall, this method keeps all possible paths open for a while, constantly adjusting the whole route as it learns more about the terrain, before finally settling on the best trail. The researchers trained their system on thousands of known protein structures, teaching it to start with a vague, uncertain idea of what the amino acids could be and then refine that idea step by step until a clear, stable sequence emerges.
The power of this new tool lies in its ability to listen to the designer. In many protein design projects, certain parts of the molecule must stay exactly the same, such as the active site where a drug binds or the disulfide bonds that hold the structure together. Other parts might just need to be generally "polar" or "charged" without a specific amino acid required. Inverse FoldDir handles both scenarios seamlessly. It can lock specific amino acids in place while redesigning the rest, or it can start with a gentle preference for certain types of amino acids in specific spots, allowing the final design to find the best fit naturally. This flexibility means the system doesn't just spit out a single answer; it explores the space of possibilities, letting different parts of the protein evolve together until they form a cohesive whole.
When the researchers tested their method against existing tools on a large set of protein shapes that the system had never seen before, it performed better than the current state-of-the-art. The designs it produced were more likely to fold back into the exact shape the scientists wanted, with a high degree of structural accuracy. The team also watched how the computer "thought" through the problem, observing that different parts of the protein settled on their final identities at different speeds. Some amino acids were decided almost immediately, while others remained flexible and changed their identity late in the process as the surrounding sequence context became clearer. This behavior mimics how natural proteins evolve, where changes in one area can ripple through the structure to stabilize the whole.
To prove that these computer-generated designs actually work in the real world, the team took on a specific challenge: redesigning an antibody fragment that binds to a green fluorescent protein. They gave the computer only the shape of the antibody, without telling it the original amino acid sequence, and asked it to create new versions that would still grab onto the target. Out of thirty-five new designs, two successfully bound to the target protein with a signal strength comparable to the original, despite having a completely different sequence. This was a significant result, showing that the computer could create a molecule that was structurally sound and functionally active, even though it looked nothing like the natural version. The success of these experiments suggests that this new approach is not just a theoretical improvement but a practical tool for creating new biological tools.
The study also highlighted what the model can and cannot do. While the designs were excellent at maintaining the physical shape of the protein, the computer's confidence in a design did not always predict whether it would bind to a target or remain stable in a cell. This is a vital distinction for future work: the method is a powerful engine for generating structurally sound candidates, but it still needs to be paired with other tools or experimental testing to ensure the protein performs its specific biological job. The researchers noted that while the system excels at stability and shape, it does not yet fully understand the complex chemistry of binding or enzymatic activity.
Ultimately, Inverse FoldDir represents a shift in how scientists approach protein design. By moving away from rigid, step-by-step generation toward a fluid, whole-sequence refinement, the method offers a more natural way to explore the vast landscape of possible proteins. It provides a way to incorporate human knowledge—like keeping a catalytic site fixed or preferring certain chemical properties—without forcing the computer into a corner. As the field moves forward, this ability to generate diverse, structurally reliable sequences with user-defined constraints could accelerate the creation of new therapies and materials, turning the abstract shapes of computer models into tangible solutions for real-world problems.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.