← Latest papers
🔬 physics

RetroDFM-R: Reasoning-Driven Retrosynthesis Prediction with Large Language Models via Reinforcement Learning

RetroDFM-R is a reinforcement learning-enhanced large language model that improves chemical retrosynthesis prediction by coupling high accuracy with transparent, step-by-step reasoning, thereby addressing limitations in generalizability and interpretability while outperforming state-of-the-art baselines on standard benchmarks.

Original authors: Situo Zhang, Hanqi Li, Lu Chen, Zihan Zhao, Xuanze Lin, Zichen Zhu, Danyu Luo, Bo Chen, Xin Chen, Kai Yu

Published 2026-08-25
📖 4 min read☕ Coffee break read

Original authors: Situo Zhang, Hanqi Li, Lu Chen, Zihan Zhao, Xuanze Lin, Zichen Zhu, Danyu Luo, Bo Chen, Xin Chen, Kai Yu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Chemists have long relied on a mental map to build new molecules, a process called retrosynthesis. Instead of starting with raw ingredients and trying to mix them into a final product, they work backward from the finished molecule, asking what simpler pieces could have been joined to create it. It is a bit like looking at a finished house and figuring out which bricks, beams, and pipes were used to build it, then tracing those materials back to their original sources. This reverse engineering is the backbone of drug discovery and materials science, allowing scientists to design efficient paths to create complex compounds. For decades, computers have tried to help with this task, but they often acted like black boxes, offering a list of ingredients without explaining the logic behind the choice. This lack of transparency made it hard for human experts to trust the machine's suggestions or to understand why a particular path was chosen over another.

A new study introduces a system called RetroDFM-R, designed to change how computers approach this chemical puzzle. Rather than simply guessing the next step, this system uses a large language model—a type of artificial intelligence trained on vast amounts of text—to reason through the problem step by step. The researchers trained the model on a massive collection of chemical reactions, teaching it not just to predict the starting materials, but to explain its thinking in plain language. The system analyzes the target molecule, identifies key structural features, and then constructs a logical argument for how to break it down, much like a human chemist would. This approach combines the ability to handle complex data with the clarity of human reasoning, making the computer's decisions visible and understandable.

The team tested this new system on a standard benchmark containing fifty thousand known chemical reactions. In these tests, the model correctly identified the starting materials for a target molecule sixty percent of the time without any extra help, and sixty-six percent of the time when using a full set of tools to explore multiple possibilities. This performance surpassed previous state-of-the-art methods, which had struggled to reach these levels of accuracy. More importantly, the system did not just get the answer right; it provided a clear, written explanation for every decision. When human experts reviewed the model's suggestions, they found that the reasoning was chemically sound and often preferred the model's proposed paths over those generated by older systems. The experts noted that the model could suggest specific reagents and reaction conditions, offering practical insights that went beyond a simple list of ingredients.

The researchers also demonstrated that the system could handle complex, real-world scenarios, including the synthesis of actual pharmaceutical drugs and specialized materials used in solar cells. In one case, the model successfully reconstructed a five-step synthesis route for a cancer drug, identifying the correct chemical transformations at each stage. In another, it mapped out the creation of a molecule designed to boost the immune system against cancer. Even when the model made a mistake, such as misidentifying a specific type of chemical bond, the error was visible in its written explanation, allowing a human to spot and correct it immediately. This transparency is a significant shift from earlier tools, which might have produced an incorrect result without any indication of where the logic failed.

To achieve this level of performance, the researchers employed a three-stage training process. First, they fed the model a curated dataset of millions of chemical examples to build a strong foundation of chemical knowledge. Next, they used a technique called distillation to teach the model how to structure its thoughts, guiding it to produce a coherent narrative before giving an answer. Finally, they used a reinforcement learning method, where the model was rewarded not just for getting the right answer, but for providing a reasoning process that was chemically logical and consistent with its final prediction. This combination of data, structured thinking, and feedback allowed the model to learn the nuances of chemical synthesis in a way that felt more like human deduction than statistical pattern matching.

The study suggests that by making the reasoning process explicit, artificial intelligence can become a more reliable partner for scientists. The system does not replace the chemist but acts as a transparent assistant that can propose routes, explain its choices, and highlight potential issues. While the model is not perfect and can still generate incorrect chemical explanations, its ability to articulate its logic makes it a powerful tool for accelerating the design of new medicines and materials. As the researchers noted, the goal is not just to predict the future of synthesis, but to make the path to that future clear and understandable for the humans who will walk it.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →