Expansion of DNA-Encoded Library Hits Using Generative Chemistry and Ultra-Large Compound Catalogs
This paper presents a synergistic methodology that combines experimentally validated DNA-encoded library (DEL) data with generative AI and ultra-large compound catalogs to rapidly expand initial hits into diverse, drug-like, and commercially available compounds, successfully identifying novel 53BP1 inhibitors with improved properties compared to the original DEL selection.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
Finding a new medicine often begins with a search through a vast, chaotic sea of chemical possibilities. Scientists need to find a single molecule that fits perfectly into a specific biological target, like a key sliding into a lock, to stop a disease process or fix a broken mechanism in the body. For decades, the standard way to do this has been high-throughput screening, a method where machines test hundreds of thousands of physical chemicals one by one. While effective, this approach is slow, expensive, and requires massive amounts of physical materials. A more modern technique, called DNA-encoded libraries, has emerged as a powerful alternative. In this method, scientists attach a unique DNA barcode to each of billions of tiny chemical building blocks. They mix these together and let them interact with a target protein. The DNA tags act as a record, allowing researchers to quickly read which chemicals stuck to the protein and which did not. This allows for the screening of enormous numbers of compounds with far less cost and effort than traditional methods. However, these libraries are limited by the specific chemicals the researchers decided to build in the first place. If the best possible drug molecule does not exist within that initial set of building blocks, the search ends there, and a potential cure remains undiscovered.
Researchers at the University of North Carolina at Chapel Hill have developed a new strategy to break through this limitation, combining the real-world data from these DNA libraries with the predictive power of artificial intelligence. Their work focuses on a protein called 53BP1, which plays a critical role in how cells repair damaged DNA. By understanding how to block this protein, scientists hope to improve gene-editing tools like CRISPR, making them more efficient. The team started with a focused collection of about 58,000 chemicals that had already been tested against the 53BP1 protein using the DNA library method. They identified a few promising candidates that showed some ability to bind to the protein. Instead of stopping there, they used these initial results to train a generative artificial intelligence model. This model does not simply look for copies of the existing chemicals; it learns the rules of what makes a molecule stick to the protein and then invents entirely new structures that follow those rules. To ensure these new ideas were practical, the team cross-referenced the AI's inventions against a massive, publicly available catalog of billions of commercially available chemicals, known as the Enamine REAL Space.
The process worked like a guided tour through a vast landscape of chemical possibilities. The researchers first took the best hits from their initial DNA library screen and used them to find a million similar molecules in the commercial catalog. They then fed this group into their AI system, which generated thousands of new, hypothetical molecules designed to bind even better to the 53BP1 protein. The system then checked these new designs against the actual shape of the protein to see if they would fit. The most promising virtual candidates were then matched back to real, purchasable chemicals in the catalog. This cycle was repeated, allowing the AI to refine its suggestions and explore even more diverse chemical shapes. The result was a list of new compounds that were structurally very different from the original ones found in the DNA library, yet they were all available to buy immediately from chemical suppliers.
When the team tested these new, AI-nominated compounds in the lab, the results were encouraging. They used a sensitive assay to measure how well the chemicals could displace a known binder from the 53BP1 protein. Three of the new compounds showed strong activity, with concentrations of 50 micromolar or less needed to achieve a significant effect, while eleven others performed well at concentrations up to 100 micromolar. These results were comparable to the best hits from the original, much smaller DNA library. Crucially, the new compounds were not just copies of the old ones; they possessed a wider variety of shapes and chemical features. They also appeared to have better properties for becoming drugs, such as higher predicted solubility in water and a structure that is easier for chemists to synthesize. The study demonstrated that by using a small amount of experimental data to guide an artificial intelligence, researchers can rapidly expand their search beyond the physical limits of their initial library. This approach allows them to find novel, potent, and commercially available molecules without the need to synthesize thousands of new compounds from scratch, offering a streamlined path to discovering the next generation of therapeutic tools.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.