PGFS++: Molecular Property Improvement under Synthesis and Diversity Constraints
This paper introduces PGFS++, a synthesis-aware reinforcement learning framework that improves molecular properties while preserving output diversity and providing explicit synthesis routes by addressing the reactant selection limitations and reward-hacking diversity collapse found in its predecessor, PGFS+.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the early stages of creating new medicines, scientists face a constant balancing act. They need to design molecules that are potent enough to fight disease, but they also need to ensure those molecules can actually be built in a laboratory. For years, artificial intelligence has helped generate new molecular structures, often optimizing them for specific traits like how well they might bind to a virus or how safe they are for the human body. However, a significant gap has remained: many of these computer-generated designs are chemically valid on paper but impossible to manufacture with the tools and ingredients available in real-world labs. To bridge this gap, researchers have turned to a method called forward synthesis, which treats drug creation like a step-by-step construction project, starting with a known molecule and adding pieces one by one using established chemical reactions. The goal is to guide the computer not just to a better molecule, but to a better molecule that comes with a clear, feasible instruction manual for how to build it.
A team of researchers at the University of Cambridge and Graphcore has developed a new system called PGFS++ to solve a specific problem that arose when trying to automate this process. Their work builds on an earlier method known as PGFS, which used a type of artificial intelligence learning called reinforcement learning. In this setup, the computer acts as an agent that tries to improve a starting molecule by selecting chemical reactions and building blocks. The previous version of this system, an upgrade they called PGFS+, worked well at finding molecules with improved properties, but it suffered from a strange flaw. The system became too efficient at finding a single, highly rewarding solution. Instead of creating a wide variety of improved versions for different starting materials, it would funnel hundreds of different input molecules into the exact same final product. The researchers call this the "magnet" effect, where the system is pulled toward one high-scoring outcome, collapsing the diversity of results and making the tool useless for creating tailored improvements for specific drugs.
To fix this, the team introduced PGFS++, a refined framework that forces the computer to keep the connection between the starting material and the final result. The system still works by selecting reaction templates and compatible building blocks from a massive library of over 118,000 available chemical ingredients. It uses a pre-computed map to quickly identify which building blocks can legally react with a specific part of the molecule, narrowing down millions of possibilities to a manageable list. The key innovation lies in how the system rewards itself. In the previous version, the computer was only rewarded if the final molecule had better properties, which encouraged it to ignore the starting point entirely. In PGFS++, the reward system includes a bonus for maintaining structural similarity to the original input. This means the computer is penalized if it strays too far from the starting molecule, ensuring that the improvements are specific to that particular drug candidate rather than a generic, one-size-fits-all solution.
The results of this new approach were tested on two common measures of drug quality: a score for how "drug-like" a molecule is and a score for how easy it is to synthesize. When the researchers ran simulations with thousands of different starting molecules, the new system successfully improved the target properties while keeping the output diverse. Unlike the earlier version, which produced nearly identical molecules for almost every input, PGFS++ generated a wide range of unique, improved structures. The system also maintained high scores for synthetic accessibility, meaning the routes it found were practical for chemists to follow. By preventing the collapse into a single "magnet" solution, the researchers demonstrated that it is possible to optimize drug candidates for specific needs without sacrificing the variety required to explore different chemical possibilities. This work suggests that future drug discovery tools can be both highly effective at improving molecular properties and respectful of the unique starting points that define real-world research.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.