How many crystal structures do you need to trust your docking results?
Analyzing 403 SARS-CoV-2 main protease crystal structures from the COVID Moonshot project reveals that structure-based pose prediction accuracy plateaus after collecting approximately five structures per generic scaffold, offering a practical guideline for optimizing resource allocation in drug discovery campaigns.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
In the high-stakes world of finding new medicines, scientists often rely on a strategy called structure-based drug discovery. Imagine trying to design a key that fits perfectly into a complex, three-dimensional lock. In this case, the lock is a protein inside the human body that causes disease, and the key is a new drug molecule. To make the key fit, researchers need to know exactly how the lock is shaped and how previous keys have sat inside it. For decades, the most reliable way to see this shape has been to freeze the protein with a drug attached and take a picture of it using X-rays, a process that creates a crystal structure. These images act as blueprints, showing scientists exactly where the drug binds so they can tweak the design of future drugs to make them stronger and more precise. However, taking these pictures is expensive, slow, and difficult. It requires growing perfect crystals and sending them to massive machines, which means scientists often face a difficult choice: how many of these expensive pictures do they actually need to take before they have enough information to move forward?
A team of researchers set out to answer this question by looking back at one of the largest and most open drug discovery efforts in history: the COVID Moonshot. This global project aimed to find a treatment for the virus that causes COVID-19 by testing thousands of different chemical designs against the virus's main protease, a critical enzyme the virus needs to survive. Over the course of the campaign, the team generated an unprecedented collection of 403 crystal structures, each showing a slightly different drug molecule bound to the enzyme. This massive dataset provided a unique opportunity to test a common assumption in the field. Many scientists believed that the more crystal structures they collected, the better their computer models would become at predicting how new, unseen drugs would fit. The researchers wanted to see if there was a point where collecting more pictures stopped helping, a moment of diminishing returns where the cost of taking another picture outweighed the benefit of the new information it provided.
To investigate this, the researchers used a method called cross-docking. Instead of just looking at the pictures they already had, they used their computers to simulate placing every single drug molecule from the collection into the binding site of every other crystal structure they had. They then checked to see if the computer could correctly predict the position of the drug based on the reference structures available. They measured success by seeing if the computer's guess was within a tiny distance of the actual, experimentally observed position. By running these simulations with different numbers of reference structures, they could watch how the accuracy improved as more data was added. They also compared two different approaches: one that used only the shape of the protein to guess the fit, and another that used the shapes of previous drugs as a guide to help the computer align the new molecules.
The results revealed a clear pattern that challenges the idea that "more is always better." The researchers found that the accuracy of the predictions depended heavily on how similar the new drug was to the drugs already in the database. When the new molecule looked very much like a previous one, the computer could predict its position with high accuracy using just a few reference structures. However, the most surprising finding was about the variety of the structures rather than the total number. The team discovered that once they had about five crystal structures for a specific chemical framework, or scaffold, adding more structures of that same type provided almost no additional benefit. The success rate of predicting the drug's position plateaued, reaching a level where it was nearly impossible to improve further by simply collecting more of the same kind of molecule.
In fact, the study showed that the diversity of the chemical structures was far more important than the sheer volume of data. For the most successful series of drugs in the campaign, having just five crystal structures of the same basic shape was enough to achieve a success rate of over 95 percent. Adding dozens more structures of that same shape did not significantly raise this number. Conversely, when the researchers tried to predict the position of a drug with a completely new chemical shape, the accuracy dropped unless they had a reference structure that matched that new shape. This suggests that the most valuable use of resources is not to keep taking pictures of slightly different versions of the same drug, but to take pictures of entirely new chemical shapes as soon as they are discovered.
The researchers also examined whether the quality of the crystal structures mattered. They looked at factors like how sharp the images were and how much the atoms seemed to wiggle in the model. They found that even with lower-quality images, the computer models could still predict the drug positions accurately, provided the chemical shapes were similar enough. This indicates that the bottleneck in the process is not the quality of the X-ray pictures, but rather the ability of the computer to generate the correct starting position for the drug. Even when the computer was given a perfect way to score the best guess, it could not improve the success rate much beyond what it achieved with a standard scoring method. This suggests that the main challenge lies in the initial generation of the drug's position, not in picking the right one from a list of possibilities.
These findings offer a practical guide for future drug discovery campaigns. Instead of spending time and money collecting hundreds of crystal structures for a single chemical series, scientists can now aim for a smaller, more strategic set. By focusing on collecting structures for a few different chemical frameworks, perhaps five to ten structures per framework, they can achieve nearly the same level of predictive power as if they had collected hundreds. This approach allows researchers to move faster and more efficiently, directing their resources toward exploring new chemical territory rather than re-examining old ground. The study confirms a long-held intuition among medicinal chemists: having a diverse set of reference points is more powerful than having a deep set of similar ones. By understanding where the returns diminish, the scientific community can make smarter decisions about how to allocate the time and money required to bring new medicines to the world.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.