← Latest papers
🔬 physics

Multi-Agent Closed-Loop Reasoning for Organic Structure Elucidation from Multimodal Spectra

The paper introduces MACROS, a multi-agent system trained on massive simulated and experimental datasets that autonomously elucidates complex organic structures from multimodal spectra with unprecedented accuracy and zero-shot generalization by emulating expert hypothesis-testing and learning fundamental chemical principles.

Original authors: Bingsen Xue, Zhuojun Jiang, Jianhao Zhang, Mingcheng Gu, Yizhe Yuan, Yongtai Zhuo, Yifan Zhang, Li Wang, Ya Su, Yue Yuan, Jiang Liu, Xueqian Kong, Cheng Jin

Published 2026-08-18
📖 7 min read🧠 Deep dive

Original authors: Bingsen Xue, Zhuojun Jiang, Jianhao Zhang, Mingcheng Gu, Yizhe Yuan, Yongtai Zhuo, Yifan Zhang, Li Wang, Ya Su, Yue Yuan, Jiang Liu, Xueqian Kong, Cheng Jin

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the laboratories where new medicines are born and where the secrets of nature are decoded, scientists face a persistent puzzle: determining the exact shape of a molecule just by looking at how it interacts with light and magnetic fields. When a chemist synthesizes a new compound or isolates a substance from a plant, they must prove what that substance actually is. They do this by running it through machines that produce complex patterns of signals, known as spectra. These signals are like a unique fingerprint for the molecule, containing clues about how its atoms are connected and arranged in space. For decades, reading these fingerprints has been a slow, labor-intensive process that requires a human expert to piece together the clues, testing one possible structure against another until the pieces fit. While computers have helped, they have struggled to replicate the flexible, logical thinking of a human expert, often failing when faced with a molecule they have never seen before or when the data is messy.

A team of researchers has now developed a new system that changes how this work is done. Instead of trying to force a computer to simply memorize a database of known molecules, they built an artificial intelligence that learns to think like a chemist. This system, called MACROS, does not just guess a structure; it proposes a candidate, checks its work, and refines its answer in a continuous loop, much like a human would. By training on millions of simulated and real-world examples, the system has learned to connect the dots between the raw signals and the three-dimensional arrangement of atoms. The result is a tool that can identify complex molecules with a speed and accuracy that surpasses current methods, even when the data is incomplete or the molecule is entirely new.

The core of this breakthrough lies in how the system is built. Rather than using a single, massive brain to do everything, the researchers created a team of specialized digital agents, each with a specific job. One agent looks at the data and suggests a possible molecular structure. Another agent takes that suggestion and simulates what the signals would look like if that structure were real. A third agent compares the simulated signals with the actual experimental data to see how well they match. If the match is poor, the system sends the information back to the first agent to try again. This cycle of proposing, testing, and refining happens automatically and repeatedly. The system learns from its own mistakes in real-time, gradually narrowing down the possibilities until it finds the structure that best explains all the observed signals. This approach mimics the iterative hypothesis-testing that human spectroscopists have used for generations, but it executes the process with a speed and consistency that humans cannot match.

What makes this system particularly remarkable is that it learned the rules of chemistry on its own. The researchers did not program it with a list of chemical rules or tell it which atoms usually connect to which. Instead, they exposed the system to over one hundred million pairs of spectra and molecular structures. Through this massive exposure, the system spontaneously discovered fundamental chemical principles. For instance, it learned that certain patterns in the data correspond to specific groups of atoms, such as rings or double bonds, without ever being explicitly told to look for them. It also developed a preference for starting its structural guesses with rigid ring systems, a strategy that mirrors the intuition of experienced human experts who know that building a molecule often begins with its most stable core. This ability to learn the underlying logic of chemistry, rather than just memorizing patterns, allows the system to generalize to molecules it has never encountered before.

The researchers tested this system on a wide variety of real-world samples, including complex natural products found in nature, synthetic compounds created in labs, and metabolites found in the human body. In many cases, the system was able to correctly identify the molecular structure using 1H and 13C nuclear magnetic resonance data, a common but often ambiguous type of signal. When the system was allowed to run through its full cycle of thinking and refining, its accuracy improved significantly. On a set of challenging natural products, the system correctly identified the structure of more than half the molecules without any prior training on those specific compounds. When the researchers fine-tuned the system for a specific type of chemistry, its success rate climbed even higher, correctly solving nearly seventy percent of the cases for metabolites. This performance held true even when the data was noisy or incomplete, demonstrating a robustness that previous computer methods lacked.

Perhaps the most compelling evidence of the system's value comes from how it works alongside human scientists. In a controlled study, professional spectroscopists were asked to solve the structures of complex molecules first on their own, and then again with the help of the system. When working alone, it took the experts about an hour to solve a single molecule, achieving a fingerprint similarity score of 0.102. When they used the system to assist them, the time dropped to just ten minutes per molecule, and the similarity score jumped to over 0.80. The system did not replace the human expert; instead, it acted as a powerful partner, handling the heavy lifting of generating and testing thousands of possibilities, allowing the human to focus on the final verification. This collaboration suggests a future where the tedious work of structure determination is automated, freeing scientists to focus on discovery and innovation.

The implications of this work extend far beyond just solving puzzles faster. By establishing a reliable method for determining molecular structures from routine data, the system opens the door to exploring the vast "dark matter" of chemistry—the millions of unknown molecules that exist in nature and in our bodies but have never been identified. Currently, many of these molecules remain invisible because the process of figuring them out is too slow and difficult. With a tool that can rapidly and accurately decode these signals, scientists could proactively discover new drugs, understand biological processes in greater detail, and accelerate the development of new materials. The system represents a shift from simply storing data to actively reasoning with it, creating a foundation for fully automated laboratories where molecular discovery happens at the speed of thought.

The success of this approach also highlights a new direction for artificial intelligence in science. Rather than trying to force scientific data into generic formats like text or images, the researchers designed their system to speak the native language of the data itself. The agents were trained directly on the raw signals and the chemical strings that represent molecules, preserving the subtle physical relationships that are often lost in translation. This specialized design allowed the system to achieve high accuracy with far fewer computing resources than general-purpose models. It suggests that the future of scientific AI lies not in bigger, more general models, but in specialized, interpretable systems that are built from the ground up to understand the specific logic of their field. By combining this specialized knowledge with a closed-loop reasoning process, the researchers have created a tool that does not just calculate, but truly understands the chemical world.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →