MaD Physics: Evaluating information seeking under constraints in physical environments
The paper introduces MaD Physics, a novel benchmark designed to evaluate AI agents' ability to infer physical laws and plan measurements under resource constraints by testing them in environments with altered physical laws, thereby revealing limitations in current Gemini models' scientific reasoning and structured exploration capabilities.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to solve a mystery, but you have a very strict rule: you only have a limited amount of money to spend on clues.
This is the core idea behind a new research project called MaD Physics (Measuring and Discovering Physics). The researchers at Google DeepMind and Mila created a video-game-like test to see how well AI "scientists" can figure out how the world works when they can't just look up the answers in a textbook and can't afford to check everything.
Here is a simple breakdown of how it works and what they found:
1. The Game: "The Mystery Box"
In this test, the AI is dropped into a virtual world where objects are moving around. But there's a catch: the rules of physics in this world are slightly broken or "altered."
- Maybe gravity doesn't pull things down the way it usually does.
- Maybe water swirls in a pattern that defies normal fluid dynamics.
- Maybe particles behave like they are connected by invisible strings that don't exist in our real world.
The AI doesn't know these rules. It has to figure them out by watching the objects move.
2. The Challenge: The "Clue Budget"
This is where the "MaD" part comes in. The AI has a budget (like a wallet with a fixed number of coins).
- Looking at an object costs money.
- Looking closely (high precision) costs a lot.
- Looking from far away (low precision) costs a little.
- Waiting longer to look costs more time.
The AI has to make smart choices: "Should I spend my whole budget checking one object very closely, or should I spend a little bit checking ten different objects to get a rough idea?"
If the AI wastes its money on bad guesses, it runs out of coins before it can solve the mystery. Once the money is gone, the game stops, and the AI must predict where the objects will be in the future based only on the clues it managed to buy.
3. The Three Worlds
To make sure the AI isn't just memorizing real-world facts, they tested it in three different "universes":
- The Bouncy Ball World (Classical Mechanics): Objects bouncing and crashing, but with weird, invisible "memory" that makes them harder to push in certain directions.
- The Swirling Water World (Fluid Mechanics): A liquid that flows, but with an invisible "alien force" pushing it in strange, spinning patterns.
- The Ghost Particle World (Quantum Mechanics): Tiny particles that act like waves, but with a twist: the rules for how they appear when you look at them are different from real life.
4. The Test Subjects
The researchers tested four different versions of Google's Gemini AI models (from a smaller, faster one to a larger, smarter one) to see how good they were at this game.
5. The Results: What the AI Got Wrong
The results were a bit humbling for the AI:
- They struggled to be "strategic": The AI often spent its budget inefficiently. It would sometimes look at the wrong things or look too closely at things that didn't matter, running out of money before it understood the pattern.
- They couldn't "invent" new laws: When the physics were altered, the AI often tried to force the real-world rules onto the new world. It was like trying to solve a Sudoku puzzle using the rules of Chess.
- Better models did better, but not perfectly: The smarter AI models (like Gemini 3) made fewer mistakes and were better at predicting the future, but they still couldn't perfectly figure out the "secret code" of the altered physics.
- Visuals made it harder: When the AI had to look at pictures of the moving objects instead of just numbers, it got even more confused.
The Bottom Line
The paper concludes that while AI is getting very good at answering questions it already knows the answers to, it is still struggling with true scientific discovery: the ability to go out, spend limited resources wisely, gather new data, and figure out rules that don't exist yet.
Think of it this way: The AI is great at being a librarian (finding existing books), but it's still learning how to be an explorer (drawing a map of a territory no one has seen before). MaD Physics is a new map for researchers to see exactly where the AI gets lost so they can help it learn to explore better.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.