What Should a Large Language Model See? Physical Invariants as a Data Representation for PDE Discovery
This paper introduces "data interpretation," a training-free method that converts raw spatiotemporal fields into theorist-relevant physical invariants as direct inputs for large language models, thereby nearly tripling the accuracy of automated partial differential equation discovery compared to using raw data.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the world of molecular science, researchers are constantly trying to understand how tiny, invisible interactions between molecules create the large-scale behaviors we can see. Imagine watching a drop of ink spread through water or a pattern of spots appear on a leopard's skin; these are macroscopic events driven by complex chemical rules. For decades, scientists have relied on human experts to write down the mathematical laws that govern these processes. This is a slow, labor-intensive task that often takes months or even years of specialized work. Meanwhile, modern experiments have become incredibly fast, capable of generating thousands of different chemical systems in a single day. The gap between the speed of data collection and the speed of human theory-building has become a major bottleneck. To bridge this, scientists have turned to artificial intelligence, hoping to automate the discovery of these governing laws. The challenge, however, is that the data these systems produce is not in a format a computer can easily "read" like a sentence. It is a vast, shifting field of values that changes over space and time, and feeding this raw information directly into an AI is often too expensive and confusing for the model to make sense of.
A new study from researchers at the California Institute of Technology proposes a clever solution to this problem. Instead of forcing the artificial intelligence to stare at the raw, overwhelming data, the team introduced a step called "data interpretation." Think of this as a translator that converts the complex, shifting field of data into a short, clear summary of physical facts—exactly the kind of notes a human expert would write down before trying to solve the puzzle. In their experiments, the researchers simulated various physical fields, such as waves or chemical reactions, and then tested whether an AI could figure out the underlying equations that created them. They compared two approaches: one where the AI was shown the raw data, and another where the AI was shown the interpreted summary. The results were striking. When the AI received the interpreted summary, it was nearly three times more accurate at recovering the correct equations than when it was shown the raw data. This improvement happened without any extra training for the AI and at a very low cost, proving that giving the machine a "theorist's perspective" is far more effective than giving it a raw data dump.
The core of this discovery lies in how the researchers transformed the data. Rather than presenting the AI with thousands of numbers representing the state of the field at every point, they first analyzed the data to extract specific physical quantities. They asked the data four simple questions: Is the system changing in a simple, smooth way, or is it oscillating like a wave? How fast are the changes happening? Is the system behaving in a straight line, or are different parts of it interacting in complex, non-linear ways? And is there a directional flow, like a current carrying something along? These questions produced a compact set of measurements that took up very little space and could be read by the AI in a fraction of a second. This process mimics how a human scientist approaches a new phenomenon: they do not memorize every single data point but instead look for the salient structures, such as shock waves or repeating patterns, that hint at the underlying rules. By feeding these insights directly to the AI, the researchers allowed the model to reason about the physics rather than just crunching numbers.
The team tested this method using a library of eight common mathematical terms that describe transport, growth, and reaction in chemical systems. They created forty-four different simulated scenarios, ranging from simple linear systems to complex ones with non-linear interactions. In the most straightforward cases, where the data was clean and the rules were simple, the AI achieved perfect accuracy when given the interpreted summary. Even in the more difficult cases involving complex, non-linear interactions, the AI significantly outperformed its attempts to guess based on raw data. The study also included a control group where the researchers shuffled the summaries, giving the AI a description of one system while asking it to solve another. In these cases, the AI's performance collapsed, dropping to the level of random guessing. This confirmed that the AI was not just memorizing common patterns but was genuinely using the provided physical measurements to construct its answers. The researchers noted that the AI struggled most when the data interpretation could identify that a complex interaction existed but could not specify exactly what form it took, highlighting that the quality of the summary is directly linked to the quality of the final answer.
This approach offers a practical path forward for automating scientific discovery. The process does not require the AI to be retrained on new data, nor does it need massive computational power. The interpretation step is fast, taking only a small fraction of a second per field, and the resulting summary is much smaller than the raw data it replaces. By separating the task of understanding the data from the task of proposing the equation, the researchers have created a pipeline where an AI can co-evolve with experimentation. As laboratories generate new data from thousands of molecular systems, this method could automatically produce candidate theories for each one, validated against the very data that created them. While the current study was limited to simulated, noise-free data, the authors suggest that with further development to handle real-world experimental noise, this technique could transform how scientists build theories, turning a process that once took years into one that happens in minutes. The work demonstrates that for artificial intelligence to truly assist in science, it may not need to see everything; it just needs to see the right things.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.