← Latest papers
🧬 biology

Transformers Discover Molecular Structure Without Graph Priors

This paper demonstrates that general-purpose Transformer models, when trained without explicit graph-based physical priors, can autonomously discover key molecular structures and interaction patterns from data alone, achieving accuracy competitive with physics-informed architectures and suggesting that such inductive biases may be learnable rather than strictly necessary.

Original authors: Tobias Kreiman, Yutong Bai, Fadi Atieh, Elizabeth Weaver, Eric Qu, Aditi S. Krishnapriyan

Published 2026-09-21
📖 5 min read🧠 Deep dive

Original authors: Tobias Kreiman, Yutong Bai, Fadi Atieh, Elizabeth Weaver, Eric Qu, Aditi S. Krishnapriyan

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Scientists have long relied on computer simulations to understand how the physical world works, from the way atoms bond to form molecules to how materials react under stress. Traditionally, these simulations are built by hand-coding the known laws of physics directly into the software. This approach ensures the results make sense, but it is often slow and expensive. In recent years, a different method has emerged: using machine learning to find patterns in data instead of following pre-written rules. While this data-driven approach has revolutionized fields like language and image recognition, its application to physics has been cautious. Many experts believed that without explicitly teaching a computer the rules of geometry and force, it would fail to understand the complex dance of atoms. The prevailing view was that a machine learning model needed a built-in map of how atoms interact to produce reliable results.

A team of researchers at the University of California, Berkeley, and Lawrence Berkeley National Laboratory set out to test the limits of this assumption. They asked a fundamental question: if you give a powerful computer model a massive amount of data about molecules but forbid it from using any pre-programmed rules about physics, can it still figure out how the physical world works on its own? To find the answer, they trained a type of artificial intelligence known as a Transformer on a vast dataset of molecular structures. This dataset, called OMol25, contains millions of examples of atoms arranged in various configurations, along with the energy and forces associated with each arrangement. Crucially, the researchers did not give the model any special instructions about how atoms should behave, nor did they force it to look at atoms in a specific geometric order. They simply fed it raw coordinates and let it learn.

The results were striking. The model, which started with no knowledge of physics, autonomously discovered the very patterns that scientists had spent decades trying to encode by hand. As the model processed the data, it learned to assign different levels of importance to the relationships between atoms based on their distance from one another. In the early stages of its learning, the model focused intensely on nearby atoms, with its attention dropping off sharply as distance increased. This behavior mirrored the way physical forces, such as electrostatic repulsion, weaken rapidly over short distances. As the model grew deeper and more complex, it shifted its focus to include longer-range interactions, eventually learning a pattern that matched the way energy decays over distance in classical physics. Remarkably, the model even figured out a natural limit to how far it needed to look to understand an atom's environment, a cutoff distance that aligned perfectly with the values human experts had manually selected for traditional models.

The researchers also observed that the model was not rigid in its approach. Instead of using a fixed rule for how far to look around an atom, it adapted its behavior based on the density of the surrounding space. In crowded areas where atoms were packed tightly together, the model concentrated on immediate neighbors. In sparse areas where atoms were isolated, it expanded its focus to include distant partners. This flexibility allowed the model to handle different molecular environments more effectively than a rigid, pre-programmed system could. Furthermore, the model learned to organize chemical elements in a way that reflected the periodic table, grouping noble gases together and ordering metals by their chemical families, all without ever being told what a periodic table was.

To ensure these findings were not just a fluke of the specific dataset, the team tested the model's ability to scale. They found that as they increased the amount of data and the size of the model, its performance improved in a predictable and consistent manner. This behavior, known as a scaling law, is a hallmark of powerful machine learning systems and suggests that the model's understanding of physics was genuine and robust. In tests where the model was asked to predict the energy of water molecules, it produced results that matched both high-level physics simulations and real-world experimental measurements. Even more surprisingly, when the researchers trained a very large version of the model, it achieved accuracy comparable to specialized models that were explicitly designed with physical laws built into their architecture.

These findings suggest that the boundary between human-engineered rules and learned knowledge is more fluid than previously thought. The study indicates that with enough data, general-purpose machine learning models can discover the fundamental principles of physical interaction on their own. This does not mean that human-designed models are obsolete; rather, it suggests that general-purpose architectures can serve as a strong starting point for scientific modeling. By relying on data to reveal structure, scientists may be able to uncover flexible solutions that are difficult to anticipate or hard to write down in advance. The work opens a new path for scientific discovery, where the computer is not just a calculator following instructions, but a learner capable of finding the rules of the universe within the data itself.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →