DiscoverPhysics: Benchmarking LLMs for Out-of-the-Box Scientific Thinking
The paper introduces DiscoverPhysics, an interactive benchmark featuring 22 simulated worlds with non-standard physics to evaluate frontier LLMs' ability to engage in long-horizon scientific reasoning by designing experiments and inferring underlying laws, revealing that while top models show promise, they still struggle with uncovering latent structures and that predictive accuracy does not guarantee conceptual understanding.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Idea: Can AI Be a Real Scientist?
Imagine you drop a smart AI into a video game universe where the rules of physics are completely made up. Maybe gravity gets weaker the farther you go, or maybe there are invisible "ghost" particles pulling things around that you can't see.
The researchers behind this paper wanted to know: Can an AI figure out the rules of this fake universe just by playing with it?
They created a benchmark called DiscoverPhysics. It's like a "final exam" for AI, but instead of asking the AI to recite facts from a textbook (like "What is Newton's Second Law?"), they ask the AI to invent the law from scratch.
How the Test Works: The "Black Box" Game
Think of the AI as a detective and the universe as a locked room with a mystery.
- The Setup: The AI is dropped into a simulated world. It can't see the "source code" (the actual math rules). It only sees particles moving around.
- The Tool: The AI has a "magic wand" (an N-body simulator). It can say, "I want to launch a particle from here with this speed," and the simulator tells it where the particle ends up.
- The Process:
- The AI launches a few particles and watches them.
- It makes a guess: "Maybe gravity works like ?"
- It tests that guess. If the particles don't move the way the guess predicts, the AI has to change its mind.
- It repeats this for many rounds, designing smarter experiments to catch the AI off guard.
- The Final Exam: The AI must submit two things:
- A Story: A plain English explanation of how the world works.
- The Code: A Python program that can predict exactly where any particle will go.
The Two Ways They Graded the AI
The researchers realized that just getting the math right isn't enough. In real science, you need to understand why things happen, not just predict them. So, they graded the AI on two things:
- The "Crystal Ball" Score (Predictive Accuracy): Does the AI's code predict where the particles will land? If the AI says "The ball will be here," and it lands there, it gets points.
- The "Teacher" Score (Conceptual Understanding): An expert human (and a second AI acting as a teacher) reads the AI's story. Did the AI actually understand the hidden rules?
- Example: If the world has invisible "dark matter" pulling things, but the AI just says "Gravity is stronger than usual," it gets a low score. It got the math right by accident, but it missed the concept.
What They Found
They tested 11 of the smartest AI models available (including top models from OpenAI, Anthropic, and open-source communities). Here is what happened:
1. The "Smartest" AIs Still Struggle
Even the best AIs only passed about 50% of the worlds. They are great at memorizing facts, but when they have to discover new rules, they get stuck.
2. The "Hidden Trap" Problem
The AIs failed most often when the universe had hidden structures.
- Analogy: Imagine you are trying to figure out why a car is slowing down. If you only look at the engine, you might guess the brakes are bad. But if there is a giant magnet under the road (hidden dark matter) pulling the car back, you won't find the answer unless you look for the magnet.
- The AIs kept trying to explain the movement with standard rules (like normal gravity) and refused to believe there were "ghost particles" or extra dimensions involved.
3. Good Math Good Understanding
One model (GPT-5.5) was amazing at predicting where particles would land (low error), but its explanation of why was often wrong. It was like a student who memorized the answer key but didn't understand the lesson.
Another model (Claude Opus) was better at figuring out the concepts (like realizing time was changing the rules), even if its math predictions were slightly less perfect.
4. The "Guessing" vs. "Experimenting" Gap
The researchers tested if the AIs were actually designing good experiments or just throwing darts at a board.
- Guided Mode: The AI chooses where to launch particles.
- Random Mode: The computer launches particles randomly, and the AI just watches.
- Result: The top AIs did much better when they could choose their own experiments. They learned to ask the right questions. However, many open-source models did just as poorly in both modes, suggesting they weren't really "thinking" about their experiments; they were just guessing.
The Verdict
The paper concludes that while AI is getting very good at solving physics problems, it is still learning how to do physics.
- Current AI: Like a brilliant student who can solve a math problem if you give them the formula, but struggles if you ask them to invent the formula based on a few messy observations.
- The Future: To be a true scientific partner, AI needs to get better at "long-horizon reasoning"—the ability to say, "My current theory is wrong, I need to try a completely different kind of experiment to find the hidden truth."
In short: The AI can predict the future, but it still needs help understanding the why behind the mystery.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.