← Latest papers
💻 bioinformatics

PharmCast: rapid generation of three-dimensional pharmacophore fingerprints from two-dimensional structure without conformer generation

PharmCast is a feedforward neural network that rapidly and accurately predicts 3D pharmacophore fingerprints directly from 2D SMILES strings without conformer generation, enabling efficient virtual screening and scaffold hopping at a fraction of the computational cost of traditional methods.

Original authors: Muskal, S. M., McGregor, M. J.

Published 2026-09-07
📖 5 min read🧠 Deep dive

Original authors: Muskal, S. M., McGregor, M. J.

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

In the world of drug discovery, scientists often think of a medicine not as a rigid object, but as a set of features waiting to be recognized. Imagine a key that does not need to look like a specific lock, but simply needs to have the right bumps and grooves to fit into the mechanism. This idea, known as the pharmacophore, suggests that a molecule's ability to bind to a biological target depends on a small collection of chemical features—like a hydrogen bond donor, an acceptor, or an aromatic ring—arranged in a specific three-dimensional pattern. For decades, researchers have used this concept to find new medicines, searching for compounds that share these spatial patterns even if their underlying chemical skeletons look completely different. This ability to spot similarities between unrelated structures, a process called scaffold hopping, is crucial for finding new treatments when the obvious chemical paths have already been explored.

However, using this three-dimensional map has always been a slow and expensive process. Because molecules are flexible and can twist into many different shapes, a computer must generate hundreds of possible 3D versions, or conformers, for every single molecule to see which features it can present. This step of generating shapes takes up nearly all the time required to analyze a compound, making it impractical to screen the millions of molecules in modern chemical libraries. The result has been that this powerful method of finding new drugs has been confined to small, carefully curated lists, unable to tackle the vast oceans of available chemistry.

A team of researchers has now developed a way to bypass this bottleneck entirely. They created a system called PharmCast, which acts as a direct translator from a molecule's flat, two-dimensional drawing to its three-dimensional potential. Instead of spending hours calculating every possible twist and turn a molecule might take, this new method uses a type of artificial intelligence to predict the final result instantly. The system was trained on nearly six million molecules, including standard drug-like compounds, active biological molecules, and flexible protein loops, learning to recognize the connection between a molecule's basic structure and the 3D features it can display.

The results show that this shortcut does not sacrifice accuracy for speed. When tested against the traditional, slow method, the new system produced nearly identical results for the vast majority of compounds. For a typical drug-like molecule, the traditional approach takes nearly three seconds to generate a single fingerprint, with the vast majority of that time spent just creating the 3D shapes. The new system completes the same task in less than a millisecond. This represents a speedup of thousands of times, turning a process that would take months to run on a massive library into one that can be finished in minutes.

The researchers tested this tool on three distinct groups of molecules to ensure it worked across different types of chemistry. They evaluated standard screening compounds, active biological molecules that had been tested in labs, and flexible protein loops that are notoriously difficult to model. In every case, the predictions matched the slow, traditional calculations with high fidelity. For the standard drug-like compounds, the error was so small it was barely noticeable, and for the flexible protein loops, the system performed even better, likely because these loops were included in the training data. The only area where the system showed slightly more variation was with large, complex biological molecules that were not well represented in the training set, a limitation the authors note is due to the specific chemistry available for learning rather than a flaw in the method itself.

One of the most significant findings is that the system successfully identifies similarities between molecules that look nothing alike on paper. In a test case involving two different HIV drugs with completely different chemical structures, the system correctly identified that they presented the same three-dimensional features to a target, a feat the traditional two-dimensional methods missed. This confirms that the tool is truly capturing the spatial essence of the molecule rather than just memorizing its chemical shape. The researchers demonstrated this capability by using the system to scan a library of over four million compounds against a reference molecule, a task that took less than half an hour. The top hits included compounds with very different structures but the same binding potential, proving the system can effectively guide the search for new chemical starting points.

The authors emphasize that while this tool is incredibly fast, it is designed to work alongside, not replace, the slower, more detailed calculations. In a real-world scenario, the fast system would be used to rank millions of candidates and narrow the list down to the most promising few. Those top candidates would then be analyzed with the traditional, slower method to confirm the exact geometry and ensure the prediction holds up. This two-step approach allows scientists to explore the full breadth of chemical space without getting bogged down by the computational cost, while still maintaining the rigorous standards required for drug development. By removing the barrier of time, this work opens the door to applying three-dimensional thinking to the entire universe of available chemistry, potentially accelerating the discovery of new medicines for diseases that currently have few options.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →