MentisOculi: Revealing the Limits of Reasoning with Mental Imagery
The paper introduces MentisOculi, a benchmark suite demonstrating that despite the potential of visual reasoning, current unified multimodal models fail to improve performance through intermediate visualizations due to compounding generation errors and an inability to effectively leverage visual aids.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to solve a complex puzzle, like a sliding block game or a paper-folding trick. You know that humans often solve these by closing their eyes and "seeing" the pieces move in their mind's eye. We call this mental imagery.
Recently, advanced AI models have started to get very good at both reading text and creating images. This led researchers to ask a big question: Can these AIs use their own "mental imagery" to help them think and solve problems, just like humans do?
To find out, the authors of this paper created a new testing ground called MENTISOCULI (which roughly translates to "Eyes of the Mind").
The Test: A Mental Gym for AI
The researchers built five different types of puzzles that are hard to describe with words but easy to solve if you can visualize them. Think of them as a gym for the brain:
- Form Board: Fitting puzzle pieces into a specific shape.
- Hinge Folding: Figuring out how to rotate connected shapes to match a target.
- Paper Fold: Predicting where holes will appear after a piece of paper is folded and punched.
- Rush Hour: Moving cars out of a traffic jam.
- Sliding Puzzle: Reassembling a scrambled picture.
These puzzles get harder, requiring more steps and more complex mental "moves."
The Experiment: Letting AI "Think" with Pictures
The researchers tested the smartest AI models available today. They gave the models a choice:
- The Old Way: Solve the puzzle using only words (like a human talking through a problem).
- The New Way: Solve the puzzle by generating intermediate images. The AI would "draw" the next step of the puzzle in its mind (or on the screen) before deciding what to do next.
It's like asking a chess player to either just think about the moves or to actually draw the board after every single move to help them plan.
The Big Discovery: Drawing Doesn't Help (Yet)
The results were surprising. Despite the models' ability to create beautiful images, generating their own "mental pictures" did not help them solve the puzzles. In fact, it often made them perform worse.
Here is why, using some simple analogies:
1. The "Hallucinating Artist" Problem
Imagine an artist who is great at drawing a single car, but when asked to draw a whole traffic jam where cars move one by one, they keep forgetting what the cars looked like in the previous frame. They might add a new car that wasn't there, make a car disappear, or change the color of the exit sign.
The paper found that when AI models tried to generate these "thinking images," they made small errors in every step. By the time they reached the final step, the image was so distorted and wrong that the AI couldn't trust it. It was like trying to navigate a maze using a map that keeps changing the walls every time you look at it.
2. The "Two Brains" Problem
The researchers discovered that the AI's "text brain" (its ability to reason with words) and its "image brain" (its ability to draw) were not talking to each other.
- The text brain could sometimes figure out the right answer if it just used logic.
- The image brain could sometimes draw a correct picture.
- But when the AI tried to use the picture to help the text brain, they got confused. The text brain would ignore the picture, or the picture would show a completely different solution than the text. They were like two people trying to steer a boat in different directions.
3. The "Oracle" Test
To see if the problem was the drawing or the understanding, the researchers gave the AI perfect, ground-truth images (like a cheat sheet showing the correct next step). Even with these perfect pictures, the AI still struggled to use them to solve the puzzle. This proved that the AI isn't just bad at drawing; it's also bad at interpreting visual information to make decisions.
The Human Comparison
The researchers also asked human students to solve these puzzles.
- Humans: When the puzzle got harder, humans spent more time thinking and took more steps to solve it. They adapted their effort.
- AI: When the puzzle got harder, the AI didn't change its strategy. It didn't spend more "thinking time" (tokens) or try harder. It just kept failing in the same way.
The Bottom Line
The paper concludes that while the idea of AI using "mental imagery" to think is very appealing, current technology isn't ready for it.
The models are like a person who can describe a car perfectly and can also draw a car perfectly, but if you ask them to draw a car driving through a traffic jam step-by-step, they lose track of the rules. They can't yet bridge the gap between generating an image and reasoning with it.
For now, these AI models are better off sticking to words for complex reasoning tasks. The "eyes of the mind" are open, but they aren't seeing clearly enough to help the brain think yet.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.