BilliardPhys-Bench: Benchmarking Physical Reasoning and Visual Dynamics of Multimodal LLMs
The paper introduces BilliardPhys-Bench, a synthetic billiards benchmark that reveals current multimodal LLMs struggle with visual dynamics and complex physical reasoning, often exhibiting a "stasis bias" where they incorrectly predict no interaction in challenging scenarios.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are watching a game of pool. You see the white cue ball get hit, and you instantly know: "It's going to hit the red ball, bounce off the side, and stop right there." Humans do this without thinking; our brains have a built-in "physics engine" that simulates the future.
This paper introduces a new test called BilliardPhys-Bench to see if AI computers have that same "gut feeling" for physics, or if they are just guessing.
Here is the breakdown of what they did and what they found, using simple analogies:
1. The Test: A Digital Pool Hall
The researchers built a virtual pool table using a computer program. They didn't use real video; they generated thousands of random scenarios where balls are hit at different speeds and angles.
- The Setup: They show the AI a single picture of the table with the balls and an arrow showing where the white ball is heading.
- The Rules: They tell the AI the laws of physics (how much friction slows the balls down, how they bounce).
- The Challenge: The AI has to predict three things without seeing the future:
- Will it hit? (Will the white ball crash into another ball?)
- Will it bounce? (Will it hit the wall?)
- Where will it stop? (Exactly where will every ball be after 1, 2, 3, 4, or 5 seconds?)
2. The Players: The AI "Students"
They tested the smartest AI models available (from families like GPT, Claude, Gemini, and Qwen). Think of these models as students taking a physics exam.
3. The Results: Who Passed and Who Failed?
The results were surprising and revealed a specific weakness in how these AIs think.
The "Stasis Bias" (The Lazy Student):
The biggest problem the researchers found is something they call "Stasis Bias." When the situation gets complicated or the prediction time gets longer (like predicting 5 seconds into the future), the AIs get scared of being wrong. Instead of guessing a collision, they default to saying, "Nothing happened." They predict the balls just sit there. It's like a student who, when unsure of the answer, writes "I don't know" instead of trying to solve the math problem.The "Good at the End, Bad at the Middle" Problem:
Some models were great at guessing where the balls would end up after 5 seconds, but terrible at explaining how they got there.- Analogy: Imagine a student who draws a perfect picture of a car crash at the end of a story but gets the whole middle part of the story wrong (saying the car drove straight when it actually swerved). They got the final state right by luck or pattern matching, but they didn't actually understand the physics of the crash.
The Winners:
The best performers were GPT-5.5 and GPT-5.4-Pro. These models were the most "balanced." They could predict the final position and correctly identify the collisions that got them there.- Interestingly, the "Pro" versions of the models (which use more computing power to "think" longer) did much better at spotting collisions than the standard versions. This suggests that giving the AI more time to "think" helps it overcome its lazy "Stasis Bias."
4. The Hard Truth
The paper concludes that while these AIs are amazing at recognizing what is in a picture (like saying "that's a pool table"), they are still very bad at predicting how things will move.
- Time is the enemy: As the prediction time got longer (from 1 second to 5 seconds), the AI's accuracy dropped. Small mistakes in the beginning added up, like a game of "Telephone" where the message gets garbled by the end.
- The Gap: There is a huge gap between "seeing" the world and "simulating" the world. Current AIs are great at describing a static photo but struggle to run a mental movie of what happens next.
Summary
The paper doesn't claim these AIs can play pool for you yet. Instead, it uses pool as a simple, clear way to show that current AI models lack a true internal model of physics. They are good at describing the present but often fail at predicting the future, especially when things get complicated or chaotic. To fix this, future AI needs to be built with better "common sense" about how objects interact, rather than just memorizing patterns.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.