← Latest papers
📄 chemistry

JURY: A Comprehensive Training-Free to Fully-Supervised Framework for Evaluating Chemical Language Models

The paper introduces JURY, a comprehensive framework utilizing four distinct probes ranging from training-free to fully supervised methods to rigorously evaluate chemical language models, revealing that their true chemical encoding capabilities are often obscured by powerful prediction heads and are more comparable to traditional fingerprints than previously thought.

Original authors: Daanish Uddin Khan, Naafey Aamer, Muhammad Sajjad, Sebastian Vollmer, Andreas Dengel, Muhammad Nabeel Asim

Published 2026-08-19
📖 6 min read🧠 Deep dive

Original authors: Daanish Uddin Khan, Naafey Aamer, Muhammad Sajjad, Sebastian Vollmer, Andreas Dengel, Muhammad Nabeel Asim

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the race to discover new medicines, scientists face a bottleneck that slows down every breakthrough: before a new chemical compound can become a drug, it must pass a rigorous series of safety checks. Researchers need to know if the body will absorb it, how it will travel through the blood, whether the liver will break it down, and if it might be toxic. These checks, known as absorption, distribution, metabolism, excretion, and toxicity, are expensive and slow to perform in a laboratory. To speed things up, scientists have turned to computers, using artificial intelligence to predict these properties just by looking at a molecule's structure. For years, the most successful tools have been "language models" trained on millions of chemical formulas written as text strings. These models learn to compress a complex molecule into a single, fixed list of numbers, a digital fingerprint that captures its chemical essence. The standard way to judge how good these models are has been to attach a simple prediction tool to them, train that tool on a small set of experimental data, and see how well it guesses the answer. If the guess is good, the model gets a high score.

However, a new study suggests that this standard method has been hiding a crucial truth. The researchers found that the prediction tool itself is often doing so much of the heavy lifting that it masks what the underlying model actually learned. A powerful tool can learn a lot just from the answers it is given, even if the digital fingerprint it is reading is weak or unhelpful. To solve this, a team of researchers from Germany developed a new way to test these models, called JURY. Instead of relying on a single trained tool, they used four different methods to read the models' digital fingerprints. Two of these methods required no training at all, simply looking at how close molecules sit to one another in the digital space. The other two methods used simple training, ranging from a basic straight-line calculation to a small, flexible network. By applying these four methods to ten different chemical models across seventy-three different safety tests, the team discovered that the models are far more similar to each other than anyone realized. No single model was the best at everything. In fact, a decades-old, hand-crafted method for describing molecules performed just as well as the most advanced modern AI in many cases.

The most surprising discovery was that the models do not differ in how much chemical knowledge they hold, but in how they organize that knowledge. Some models arrange their digital fingerprints so that molecules with similar properties naturally sit close together, making them easy to find without any extra training. Other models hide that same chemical knowledge in a way that is only visible after a predictor is trained to rearrange the data. It is like looking at a library where one librarian has already sorted books by genre, while another has stacked them randomly but knows exactly where to find a specific book if you ask the right question. The study showed that a model that excels at the first approach might fail at the second, and vice versa. This means that a single score cannot tell you if a model is good; you must look at its profile across different testing methods to understand where its strengths lie.

The researchers tested ten different chemical language models, including several that use three different transformer families, alongside a traditional hand-crafted method known as an extended-connectivity fingerprint. They applied their four testing methods to seventy-three different tasks, covering everything from how a drug is absorbed to how toxic it might be. When they looked at the results, they found that the ranking of the models changed completely depending on which testing method was used. A model that came in first place when tested with a method that required no training would often drop to the bottom of the list when tested with a method that involved training a predictor. Conversely, a model that struggled without training would rise to the top once a simple predictor was added. This inconsistency proved that the models were not simply "better" or "worse," but that they stored their chemical knowledge in fundamentally different ways.

One model, called MoLFormer-XL, demonstrated that it had arranged its digital space so clearly that molecules with similar behaviors were already grouped together. It performed exceptionally well when the researchers simply looked at the nearest neighbors in the digital space, without training any extra tools. In contrast, another model, MolEncoder, performed poorly when the researchers looked at the raw arrangement of the space. However, once a simple predictor was trained on its data, MolEncoder jumped to the top of the rankings. This showed that MolEncoder held the same amount of chemical knowledge, but it was hidden in a format that required a trained tool to unlock. The study also found that the traditional, hand-crafted fingerprint method, which predates modern AI, often outperformed the most advanced models when no training was allowed, particularly in predicting how chemicals behave in the body. This suggests that for certain tasks, the old way of organizing chemical data is still highly effective.

The team concluded that the field has been relying on a single number to judge complex systems, which is like judging a musician by only one song. A model might be excellent at one type of prediction but terrible at another, depending on how its internal data is structured. The new framework, JURY, replaces that single score with a detailed profile that shows how a model performs across different levels of supervision. This allows scientists to choose the right tool for the job. If a researcher needs to quickly screen a library of chemicals without spending time training a new tool, they should choose a model that performs well on the untrained tests. If they have plenty of data and want to train a specific predictor for a high-stakes task, they might choose a model that excels only after training. By understanding where a model's knowledge sits and how it is organized, scientists can make better decisions about which tools to use, moving beyond simple rankings to a deeper understanding of what these artificial intelligences have actually learned.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →