LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?
The paper introduces LEGO-Puzzles, a benchmark evaluating the multi-step spatial reasoning capabilities of Multimodal Large Language Models (MLLMs) through elementary tasks and assembly planning, revealing that even the strongest models significantly underperform compared to humans and struggle to maintain accuracy as planning complexity increases.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a robot how to build a complex castle out of blocks, but you can only talk to it and show it pictures. This is the world of Multimodal Large Language Models (MLLMs). Think of these models as super-smart digital brains that can read text and "see" images, trying to understand the world just like humans do. One of the trickiest things for any brain to learn is spatial reasoning—the ability to understand how objects fit together in 3D space, knowing which block goes on top, which one needs to be turned around, and how the whole structure changes as you add pieces. While these AI models are getting better at chatting and describing pictures, scientists are still wondering: Can they actually plan a sequence of moves to build something? It's one thing to look at a finished tower and say, "That's tall," but it's a whole different challenge to figure out the exact order of 50 steps to build it from scratch. If we want robots to fix our cars or help us assemble furniture, they need to master this kind of step-by-step thinking.
Enter LEGO-Puzzles, a new "exam" designed by researchers to test exactly this skill. Instead of using real plastic bricks, the researchers created a massive digital playground using LEGO instructions to see how well the smartest AI models in the world can handle spatial puzzles. They broke the test down into two main levels. The first level, the Elementary Set, is like a warm-up. It asks the AI simple questions about a single picture: "Which block is taller?" "Are these two pieces touching?" or "If I turn this piece, what does it look like from the side?" This tests basic "spatial understanding." The second level, the Planning Set, is the real boss battle. Here, the AI isn't just answering questions; it has to act like a master builder. It is given a starting pile of blocks and a picture of the final goal, and it must generate a step-by-step plan to get there. The researchers tested 29 of the most advanced AI models available, including the biggest names like GPT-5 and Gemini, and even asked human experts to take the same test.
The results were a bit of a reality check for the AI world. Even the strongest models struggled with the basics. In the elementary questions, the best AI models fell at least 20% behind human performance. For example, when asked to judge the height of blocks in 3D space, many models got confused, looking at the flat 2D picture on the screen instead of imagining the true 3D shape. But the trouble got much worse as the tasks got longer. When the AI had to plan a sequence of steps using multiple-choice questions, their accuracy started to drop significantly once the plan exceeded 3 steps, falling below 50%. By the time the plan required 8 steps, the accuracy of the best models dropped to 0%. In contrast, human participants solved every single planning task perfectly, no matter how long the chain of steps was.
The researchers also tried something even harder: asking the AI to not just describe the steps, but to actually draw the intermediate pictures of the building process. They wanted to see if the AI could imagine what the castle looks like after step 1, then step 2, and so on. The results were stark. Even for a short plan of just 3 steps, the models failed completely at generating the correct visual sequence. They could sometimes draw a picture that looked like a LEGO brick, but they couldn't place it in the right spot or get the orientation correct. The paper notes that because the models failed so badly at just 3 steps for image generation, the researchers didn't even test longer horizons for this specific task. This suggests that while these models are great at recognizing patterns, they lack the ability to "imagine" the consequences of their actions over time. They can't hold a mental map of the building process in their heads without losing track.
In short, LEGO-Puzzles reveals that while our current AI models are impressive, they are still very far from having true spatial intelligence. They can't reliably plan a multi-step assembly, and they fail miserably when asked to visualize the future states of an object. The paper concludes that we have a long way to go before AI can truly understand the physical world well enough to help us build things, fix broken items, or navigate complex environments on its own. The gap between human intuition and machine calculation in the physical world is still wide, and this new benchmark gives us a clear map of exactly where the models are stumbling.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.