Benchmarking antibody-antigen co-folding on human monomeric antigens
This study introduces the HuMonoAg-Bench benchmark to evaluate ten antibody-antigen co-folding protocols, revealing that while recent methods have significantly improved CDRH3 modeling and achieved medium-or-better accuracy for about half of post-cutoff complexes, challenges persist in sampling correct binding modes and modeling flexible loops, limiting the prediction of many structurally heterogeneous complexes.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
In the microscopic world of the human immune system, antibodies act as highly specialized keys, designed to fit perfectly into the locks presented by invading viruses or harmful proteins. For decades, scientists have dreamed of being able to predict exactly how these keys and locks fit together just by looking at their chemical blueprints. This ability would revolutionize medicine, allowing researchers to design new drugs and vaccines without needing to build and test every single version in a lab. However, this prediction has remained one of the most stubborn puzzles in biology. While computers have become excellent at guessing how two proteins might stick together, antibodies are a special case. Unlike most proteins that evolve alongside their partners, antibodies evolve independently, meaning they lack the shared evolutionary history that helps computers make accurate guesses. Furthermore, the part of the antibody that actually grabs onto the target is a flexible, wiggly loop that changes shape constantly, making it incredibly difficult to model with standard tools.
A team of researchers set out to measure how far the latest generation of computer programs has come in solving this specific puzzle. They gathered a massive collection of 412 real-world examples of antibodies bound to human proteins, creating a rigorous test set that had not been seen by the computer programs before they were built. By running ten different prediction methods against this new data, they discovered that the most recent software has made a dramatic leap forward, successfully predicting the correct shape for about half of the new, unseen cases. However, the study also revealed that even the best programs still struggle with the most flexible parts of the antibody and that simply guessing more times does not always lead to a better answer. The researchers found that if they could tell the computer exactly where the antibody should touch the target, the older programs could catch up to the newer ones, suggesting that the main hurdle is not just the software's intelligence, but its ability to find the correct starting point in a vast sea of possibilities.
The researchers began by assembling a benchmark called HuMonoAg-Bench, a curated library of 412 antibody complexes involving human proteins. They carefully split this collection into two groups based on when the structures were discovered. One group consisted of 278 complexes released before September 2021, a date that serves as a cutoff for the training data of many modern prediction tools. The other group contained 134 complexes released after that date, ensuring that the computer programs had never seen these specific shapes before. To make the test even more precise, they further divided the new group into those involving antigens that were very similar to older ones and those that were completely novel. This setup allowed them to test whether the computers were simply memorizing old examples or truly learning how to predict new interactions. They then ran ten different co-folding protocols—methods that simulate how two proteins fold together—against this dataset, comparing the results to the actual experimental structures to see how close the predictions were.
The results showed a clear and significant improvement in the newest methods compared to older ones. When predicting the structures of the new, unseen complexes, the older software managed to find a correct or near-correct shape in only about 16 to 22 percent of cases. In contrast, the most advanced methods, such as Protenix v2, IntelliFold-2, and ESMFold2, succeeded in finding a medium-or-better quality model for roughly 33 to 52 percent of the same cases. The best performer, Protenix v2, achieved a success rate of 52 percent for standard antibodies and an even higher 82 percent for a specific type of single-domain antibody when the antigen was familiar. Despite this progress, the study emphasized that nearly half of the new targets remained unsolved, indicating that while the field has advanced rapidly, a large fraction of antibody interactions are still beyond the reach of current technology. The researchers noted that this improvement was not simply because the newer programs had seen more data, as they shared similar training cutoffs, but rather due to fundamental improvements in how they model the complex shapes.
A deeper look into why some predictions succeeded and others failed pointed to a specific structural feature: the CDRH3 loop. This is a flexible, hairpin-like structure on the antibody that acts as the primary point of contact with the target. The study found that the accuracy of this single loop was the strongest indicator of whether a prediction would be successful. While the computers were generally good at modeling the rigid parts of the antigen and the other, more stable loops of the antibody, the CDRH3 loop remained the most difficult component to get right, especially when it was long. The newer methods showed a marked improvement in modeling this specific loop, which directly correlated with their higher overall success rates. However, even when the local structure of the loop was modeled well, it did not guarantee that the entire antibody would be positioned correctly relative to the target, showing that getting the pieces right is not the same as assembling the whole puzzle correctly.
The researchers also tested what would happen if they gave the computer programs a hint about where the antibody should bind. In a real-world scenario, scientists might know a few key spots on the target protein where an antibody attaches, but they do not know the exact shape of the antibody. By supplying this limited information as a constraint, the older methods saw their success rates jump by 20 to 30 percentage points, bringing them up to the level of the strongest unconstrained methods. This finding suggests that the primary limitation for the older software was not a lack of understanding of the physics, but an inability to find the correct binding mode among billions of possibilities. When the search space was narrowed by a hint, the older programs could find the solution. However, this benefit depended heavily on the accuracy of the hint; if the provided information was only partially correct, the improvement vanished.
Finally, the team investigated whether running the simulations more times would solve the remaining problems. They found that for the unconstrained methods, the main issue was sampling failure: the correct shape simply never appeared in the set of generated models. When they increased the number of simulation runs, they did find more correct shapes, but they also found that the computer's ranking system often failed to pick the best one as the top answer. This created a new bottleneck: even when the correct answer was generated, the software did not always recognize it as the best option. The study concluded that while increasing the number of attempts helps, the ability to correctly rank and select the best model is becoming just as important as the ability to generate them. The remaining unsolved cases were a diverse mix of difficult targets with no single common feature, suggesting that the path forward requires solving multiple different types of challenges rather than fixing just one specific problem.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.