← Latest papers
🤖 AI

Do Vision-Language-Models show human-like logical problem-solving capability in point and click puzzle games?

This paper introduces VLATIM, a benchmark based on *The Incredible Machine 2* that reveals a significant gap between the high-level planning abilities and precise visual grounding of current Vision-Language Models, demonstrating that they have not yet achieved human-like logical problem-solving capabilities in interactive point-and-click puzzle environments.

Original authors: Dominik Helfenstein, Marco Menner, Maximilian Triebel

Published 2026-05-13
📖 4 min read☕ Coffee break read

Original authors: Dominik Helfenstein, Marco Menner, Maximilian Triebel

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a very smart robot assistant. You can talk to it, show it pictures, and ask it to solve a puzzle. The big question this paper asks is: Is this robot smart enough to think and act like a human when playing a tricky "point-and-click" puzzle game?

To find out, the researchers created a special test called VLATIM. They used a classic game from the 90s called The Incredible Machine 2. In this game, you don't just press buttons; you have to build crazy Rube Goldberg-style machines. For example, you might need to drop a bowling ball on a seesaw to launch a cat, which scares a mouse, which turns on a generator to power a mixer. It requires understanding physics, cause-and-effect, and precise mouse movements.

Here is how the paper breaks down the robot's performance, using simple analogies:

1. The Five-Step Test

The researchers didn't just throw the robot into the deep end. They built a ladder with five rungs, getting harder each time:

  • Rung 1 (Spotting Things): Can the robot point its finger at a specific object on the screen? (e.g., "Where is the candle?")
  • Rung 2 (Knowing the Rules): Can the robot answer questions about how things work? (e.g., "Which wall is slippery?")
  • Rung 3 (Predicting the Future): Can the robot guess what happens next? (e.g., "If I drop the ball, where will it land?")
  • Rung 4 (Moving Things): Can the robot actually use the mouse to drag, drop, and rotate objects?
  • Rung 5 (Solving the Whole Puzzle): Can the robot look at a blank screen, figure out the goal, and build the whole machine from scratch?

2. The Two Types of "Robots" Tested

They tested five different AI models, which fell into two distinct personality types:

  • The "Blind Strategists" (Big, expensive models like GPT and Gemini):

    • The Metaphor: Imagine a brilliant chess grandmaster who is wearing thick, blurry foggy glasses.
    • What they did: They were amazing at the thinking parts. They understood the physics, made great plans, and knew exactly what needed to be done.
    • The Failure: When it came time to actually click the mouse, they were terrible. They couldn't see exactly where the object was. They would plan a perfect move but click 10 inches to the left, missing the target completely. They were "blind" to the precise details needed to execute their smart ideas.
  • The "Myopic Operators" (Specialized models like UI-Tars):

    • The Metaphor: Imagine a very precise surgeon with steady hands, but who has no idea what surgery they are supposed to perform.
    • What they did: They were great at clicking the right spot. If you told them "click here," they did it perfectly.
    • The Failure: They couldn't think ahead. They didn't understand the chain reaction. They would click the right button but in the wrong order, or get stuck in a loop doing the same thing over and over, never realizing they were failing.

3. The Verdict: The Gap Between Thinking and Doing

The paper's main conclusion is that current AI models do not yet have human-like problem-solving skills in these games.

  • The Disconnect: There is a huge gap between planning and doing. The smartest models can plan a solution but fail to execute it because they can't "see" precisely enough. The models that can "see" precisely aren't smart enough to plan the solution.
  • The Result: In the final test (solving the whole puzzle), every single model failed. None of them could successfully build the machine from start to finish.

4. Why This Matters (According to the Paper)

The researchers wanted to see if AI could just "read the manual" and play the game like a human. They found that while AI is getting better at understanding language and images separately, it still struggles to combine logical reasoning (thinking) with continuous action (precise mouse control).

In short: The AI is like a genius architect who can draw a perfect blueprint but cannot hold a hammer steady enough to build the house. Until the AI can do both at the same time, it cannot truly solve these puzzles like a human.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →