← Latest papers
💻 bioinformatics

Targeted finetuning enables co-folding models to learn ligand-induced protein conformational states

This paper demonstrates that targeted finetuning of foundation models like Boltz-1 can overcome training data biases to accurately predict novel ligand-induced protein conformational states and allosteric binding sites, thereby enabling more effective drug discovery applications.

Original authors: Gorantla, R., Schleberger, C., Sesterhenn, F.

Published 2026-09-27
📖 3 min read☕ Coffee break read

Original authors: Gorantla, R., Schleberger, C., Sesterhenn, F.

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

For decades, scientists have sought a way to predict the three-dimensional shape of a protein just by looking at its genetic code. This is no small feat, as proteins are not rigid rods but flexible chains that twist and turn into complex shapes to perform their jobs. In recent years, powerful computer models have emerged that can predict these shapes with remarkable accuracy, even when a protein is bound to a small molecule, like a drug. These models, often called co-folding tools, are trained on vast libraries of known structures, learning the general rules of how proteins and molecules fit together. However, a significant hurdle remains: these models often struggle when faced with a protein that changes its shape in response to a new drug, or when the drug binds to a part of the protein that was never seen before. If a model cannot see these changes, it cannot help researchers design medicines that rely on them, leaving a gap between what the computer predicts and what actually happens in a living cell.

A team of researchers recently addressed this gap by showing that the failure of these models is not due to a flaw in their design, but rather a bias in the data they were taught. They demonstrated that by carefully teaching the model a few specific new examples, it could learn to predict complex, unseen behaviors. Using ten new X-ray images of a protein called Werner helicase, which is involved in DNA repair and was studied during a drug discovery program, the researchers took a powerful existing model and gave it a targeted lesson. They showed it how this protein locks itself into an inactive state when a specific drug binds to an unusual spot, a change that the original model had never encountered. The result was a model that could not only predict this new, locked shape but also correctly identify the drug's binding site, all while keeping its original ability to predict the protein's normal, active shape.

The researchers found that this new, trained model did more than just memorize these ten specific images. It learned the underlying logic of how the drug forces the protein to change, allowing it to generalize to different chemical series of drugs that it had never seen before. Furthermore, this new understanding transferred to other proteins in the same family, known as RecQ helicases, but only when the specific sequence of the binding site matched the logic the model had learned. This suggests that the model had truly grasped the mechanism of the change rather than just copying the shapes. The work indicates that these powerful computer models can be adapted as new structural data becomes available, allowing them to capture the subtle, ligand-induced switches that are central to how biology is regulated and how medicines might intervene. By showing that the limitation lies in the training data rather than the architecture, the study provides a clear path forward for updating these tools to handle the dynamic, shifting nature of proteins in the real world.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →