← Latest papers
🧬 biology

An energy-based model empowers ultra-large virtual screening against 100-billion-scale chemical libraries

The paper introduces DrugJEPA, an energy-based model that enables ultra-fast and cost-effective virtual screening of 100-billion-scale chemical libraries with state-of-the-art accuracy, successfully identifying potent nanomolar hits across diverse protein targets in wet-laboratory validation.

Original authors: Shengyong Yang, Yuanyuan Jiang, Chong Huang, Mengzhe Dai, Rui Yao, Weining Sun, Liting Zhang, Qiao Huang

Published 2026-09-07
📖 4 min read☕ Coffee break read

Original authors: Shengyong Yang, Yuanyuan Jiang, Chong Huang, Mengzhe Dai, Rui Yao, Weining Sun, Liting Zhang, Qiao Huang

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

The search for new medicines often begins with a needle in a haystack, but in this case, the haystack contains nearly one hundred billion different chemical structures. For decades, scientists have relied on two main ways to find the right needle: physically testing millions of compounds in a lab, or using computers to simulate how they might fit into a biological target. Physical testing is slow and expensive, while traditional computer simulations, though faster, are still too sluggish to handle libraries of this immense size. The challenge has been to find a method that is both fast enough to scan a hundred billion molecules and accurate enough to spot the few that actually work, without requiring a supercomputer the size of a city.

A team of researchers at Sichuan University has developed a new approach that bridges this gap. They created a sophisticated computer model called DrugJEPA, which acts like a highly trained intuition engine for drug discovery. Instead of trying to simulate every single atomic movement, which takes too long, the model learns to recognize the "shape" of a good match by studying millions of examples of how proteins and molecules interact. It uses a technique that allows it to predict the relationship between a protein's binding pocket and a drug molecule, effectively filtering out the billions of useless candidates in a matter of hours. When tested against three very different types of biological targets, the system didn't just find hits; it found potent ones, including some that work at incredibly low concentrations.

The core of this achievement lies in how the researchers built their library and their model. They started by gathering a massive collection of potential drug candidates from existing databases, converting them into three-dimensional shapes, and storing them in a system they named ZEUS-3D. This library holds roughly ninety-six billion unique molecules, with hundreds of billions of different 3D variations to account for how they might twist and turn in a solution. To search this ocean of data, they needed a tool that could move at lightning speed. They built DrugJEPA, a neural network that learns to compress the complex information of a protein and a molecule into a simple numerical code, or "embedding." By comparing these codes, the computer can instantly tell how well a molecule might fit a protein, a process that is millions of times faster than traditional methods.

To prove their system worked, the team ran a massive virtual screening campaign. They set the model loose on their hundred-billion-molecule library to hunt for drugs against three distinct biological targets: a kinase involved in immune response, a receptor in the brain linked to mental health, and an enzyme that modifies RNA, for which no known small-molecule drugs existed. The entire search, covering a quadrillion potential interactions, was completed in just twenty-eight hours using only four standard graphics cards. This speed represents a leap forward, making it possible for ordinary drug development teams to explore chemical spaces that were previously accessible only to the largest corporations with massive computing resources.

The results of the search were then put to the test in the real world. The researchers took the top candidates identified by the computer and synthesized or purchased them to see if they actually worked in a laboratory setting. For the immune-related target, the model identified active compounds with a success rate of forty percent, and several of these molecules were powerful enough to stop the target's activity at concentrations as low as one nanomolar. For the brain receptor, the hit rate was twenty percent, and the system even rediscovered a known drug that had been overlooked for this specific target, confirming the model's ability to find both new and existing solutions. Most impressively, for the RNA-modifying enzyme, which had no known drug candidates, the model found a thirty percent hit rate and identified a molecule that bound tightly to the protein, effectively blocking its function.

To ensure these findings were not just lucky guesses, the team looked deeper into how the best candidates interacted with their targets. For the immune target, they determined the exact 3D structure of the drug bound to the protein, confirming that it sat precisely where the computer predicted it would. For the brain receptor and the RNA enzyme, where creating a physical crystal structure proved difficult, they used targeted mutations to change specific parts of the protein. When they altered the parts the model predicted were important for binding, the drugs stopped working, providing strong evidence that the model had correctly identified the mechanism of action. This combination of speed, scale, and experimental validation suggests that the barrier to exploring the vast universe of possible medicines has been significantly lowered, offering a practical path to finding new treatments for diseases that have long resisted discovery.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →