← Latest papers
🤖 AI

Library Reachability in LSR-Synth: How Anti-Memorization Design Changes the Measurement of Symbolic Discovery

This paper evaluates the LSR-Synth benchmark's ability to distinguish between scientific prior knowledge and conventional operator search by demonstrating that while its anti-memorization controls are effective, current tasks primarily measure the recombination of unseen expressions within a fixed vocabulary rather than the unique contributions of language model priors unless that vocabulary is deliberately constrained.

Original authors: Zhan'ao Yao, Liang Yin, Zhihao Gao, Boxuan Zhang, Xiaoyu Wu, Linjing Li, Rongyan Wang, Tingwei Chen, Youwei Wang, Xiaolin Zhao, Jiahui Shi, Jianjun Liu

Published 2026-08-03
📖 3 min read☕ Coffee break read

Original authors: Zhan'ao Yao, Liang Yin, Zhihao Gao, Boxuan Zhang, Xiaoyu Wu, Linjing Li, Rongyan Wang, Tingwei Chen, Youwei Wang, Xiaolin Zhao, Jiahui Shi, Jianjun Liu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to solve a mystery: Is a brilliant new detective actually using their sharp intuition to crack a case, or are they just reciting a solution they memorized from a famous book? This is the big question facing scientists who use Artificial Intelligence (AI) to discover new mathematical laws. For a long time, researchers have been training AI to look at data—like the speed of a falling apple or the growth of bacteria—and guess the hidden formula that explains it. This field is called "symbolic regression." The problem is that AI models are like giant libraries that have read almost everything ever written. If you ask them for the formula for gravity, they might just be remembering it from a textbook they read during training, not actually figuring it out from the data. To fix this, scientists created a special test called LSR-Synth. They invented brand-new, made-up scientific problems by mixing known rules with strange, new ingredients. The hope was that since these problems were totally new, the AI couldn't just "remember" the answer; it would have to truly discover it.

But here is the twist: even if the final answer is new, the tools the AI uses to find it might be old. Imagine a chef trying to invent a new dish. If the kitchen is stocked with every spice and ingredient imaginable (a "fixed library"), the chef might not need to be a genius to make a great meal; they just need to mix the right existing ingredients. This paper asks a very specific, slightly boring but crucial question: Do these new AI tests actually prove the AI is using its "brain" (scientific knowledge), or is it just a really good mixer of ingredients that were already in the kitchen? The authors built a "blind" kitchen where the AI couldn't see the names of the ingredients or the type of food, only the raw numbers. They compared a smart AI chef against a simple, rule-following robot that just tries every possible mix of standard ingredients.

The results are a bit of a reality check for the hype. The authors found that in most of these "new" scientific puzzles, the simple robot with the standard kitchen tools could solve them just as well as the smart AI chef. In fact, when they gave the AI chef a full kitchen, it didn't solve many more puzzles than the robot did. The AI only showed a real advantage when the kitchen was stripped down, leaving out essential ingredients like "exponentials" or "trigonometric functions." Only then did the AI's ability to suggest new, weird ingredients help solve the problem.

So, what does this mean? It doesn't mean the AI is useless or that it is just using a memorized solution. It means that for the current batch of tests, the "standard kitchen" is so well-stocked that the AI's special "scientific intuition" isn't strictly necessary to get a high score. The paper suggests that to truly prove an AI is discovering new science and not just rearranging old tools, we need to design tests where the standard tools fail first. Until then, a high score on these tests might just mean the AI is good at using a standard toolbox, not that it has unlocked a new kind of scientific genius. The authors conclude that while these tests are great at stopping AI from just copying old formulas, they aren't quite good enough yet to prove the AI is adding something truly new from its own mind.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →