MolGlueBench reveals heterogeneous transfer of context gains across molecular-glue DC50 domains
MolGlueBench introduces a domain-aware benchmarking framework that reveals how internal scaffold generalization gains in molecular-glue potency prediction do not consistently transfer across fixed database and target domains, highlighting the risks of overstating model performance when chemical and experimental contexts recur.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
In the quest to cure diseases, scientists often look for tiny molecules that can act as molecular glue. These are small chemical compounds designed to stick two specific proteins together inside a cell. When these two proteins are forced to interact, the cell's natural cleanup crew is often tricked into destroying a harmful target protein, such as one that drives cancer growth. This approach offers hope for treating diseases that were previously considered untreatable by traditional drugs. However, finding the right glue is incredibly difficult. Scientists must test thousands of chemical structures to see which ones successfully trigger this destruction, and the process is slow, expensive, and often unpredictable. To speed things up, researchers have turned to computers, hoping to build models that can predict how well a new molecule will work before it is ever synthesized in a lab.
The promise of these computer models is that they can learn from past experiments to forecast future success. But there is a hidden trap in how these models are usually tested. Often, a computer program is shown a set of past data, asked to learn the patterns, and then tested on a different slice of that same data. If the test data shares the same experimental conditions or comes from the same source as the training data, the model might appear to be a genius. In reality, it may have simply memorized the quirks of the specific laboratory or database it was trained on, rather than learning the true rules of chemistry. This creates a false sense of confidence, where a model looks perfect on paper but fails completely when faced with a new type of experiment or a new biological target.
A team of researchers at Nantong University and the National University of Singapore set out to expose this problem with a new, rigorous test called MolGlueBench. Instead of just checking if a model can guess the right answer for a familiar molecule, they designed a challenge that forces the model to prove it can handle completely new situations. They gathered 1,560 specific measurements of how well different molecular glues worked, drawn from four different public databases. These measurements covered over a thousand unique chemical compounds and hundreds of different structural frameworks. Crucially, they recorded not just the chemical structure of the glue, but also the full context of the experiment: which cell line was used, which proteins were involved, and which database the data came from.
The researchers then put their models through a series of increasingly difficult tests. First, they asked the models to predict the performance of new chemical structures that had never been seen before, but which came from the same mix of experiments as the training data. In this scenario, the models performed reasonably well. When the computer was allowed to use both the chemical structure and the experimental context, it improved its predictions slightly, suggesting that knowing the conditions of the test helped the model rank the molecules correctly. This result looked promising, as if the extra information about the experiment was a valuable clue.
However, the story changed dramatically when the researchers changed the rules to simulate a real-world scenario. They asked the models to predict the performance of molecules in entirely new databases or for entirely new target proteins that the model had never encountered during training. This is the true test of whether a model has learned a universal rule or just memorized a specific dataset. In these new, unseen environments, the advantage of using experimental context vanished. In fact, adding the extra context information often made the predictions worse. The models that relied on the experimental details failed to transfer their knowledge to the new domains, while the models that relied solely on the chemical structure performed slightly better, though still with significant errors.
The study revealed that the helpfulness of experimental context is not a universal truth but a local shortcut. Inside the familiar data, the context helped the model guess the right order of molecules. But when the model stepped outside its comfort zone into a new database or a new biological target, those same context clues became misleading. The model had learned to associate certain experimental setups with certain outcomes, but those associations did not hold up when the setup changed. For example, the models struggled particularly when trying to predict results for specific targets like GSPT1 and CDK2, or when moving away from data sourced from one specific database called TPDdb. The errors were not small; the models' predictions were off by a factor of roughly eight to ten times the actual concentration needed to degrade the target protein.
This finding serves as a crucial reality check for the field of drug discovery. It demonstrates that a model's ability to predict well within a single collection of data does not guarantee it will work for future discoveries or for different types of biological problems. The researchers concluded that while these computer models can help rank molecules within a known set of experiments, they cannot yet be trusted to predict absolute values or to prioritize candidates for new, unseen targets. The study does not say that context is useless, but rather that its value is fragile and depends entirely on the specific conditions of the data. Until models can prove they work across these different domains, scientists must remain cautious about relying on them to make final decisions about which molecules to test in the lab. The path forward requires building models that are tested against these harder, more realistic challenges, ensuring that the tools we build are truly robust enough to handle the complexity of the living world.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.