JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles
The paper introduces JigShape, a jigsaw puzzle benchmark with interlocking pieces that reveals current vision-language models lack robust geometric reasoning, as they fail to scale beyond small grids despite strong performance on simple 4×4 puzzles.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a robot how to understand the world, not just by recognizing what things look like, but by understanding how they fit together in space. This is the realm of Vision-Language Models (VLMs), a type of artificial intelligence that can "see" images and "read" text, often chatting about what it sees. While these robots are getting incredibly good at naming objects or describing scenes, scientists are starting to wonder: Can they actually reason about space? Can they look at a scattered pile of objects and figure out exactly where each one belongs, just like a human would? This isn't just a party trick; it's a fundamental skill needed for robots to pick up tools, for architects to design buildings, and for any AI to truly understand the physical world around it. If an AI can't figure out how pieces fit together, it's like having a librarian who can read every book but can't organize the shelves.
Enter a new study called JigShape, which decides to test these AI brains with a classic childhood challenge: the jigsaw puzzle. But not just any puzzle. The researchers realized that previous tests were a bit like cheating. They used puzzles where the pieces were just simple squares cut from a photo. If you had a picture of a blue sky, any blue square could fit anywhere, making it impossible to tell if the AI was actually "thinking" or just guessing. To fix this, the team created a new benchmark using tab-and-blank pieces—the kind with the little bumps and holes that lock together. This adds a strict rule: a bump must go into a hole. It turns the puzzle from a vague guessing game into a precise logic problem where there is only one correct answer.
The researchers built a massive library of these puzzles, ranging from small 4x4 grids to huge 16x16 grids with hundreds of pieces. They then asked the world's smartest AI models to solve them. The results were a mix of surprise and disappointment. When the puzzles were small (4x4), one top-tier model, GPT-5.5, actually did quite well, achieving a Piece Accuracy of 69.65% (meaning that, on average, nearly 70% of the individual pieces were placed in the correct spots). It seemed to understand both the picture and the shape of the pieces. However, almost every other model, even those designed to "reason" better, failed miserably, performing no better than if they had just picked pieces at random.
The real story, though, is what happened when the puzzles got bigger. As the grid grew to 8x8 and 12x12, the performance of even the best models crashed. The model that was doing great on small puzzles suddenly dropped to near-random guessing on larger ones. The researchers call this a "scaling cliff." It suggests that while these AIs can handle a few pieces, they completely lose their ability to keep track of all the rules when the number of pieces gets high. They can't maintain the "lock-and-key" logic across a large board.
The team also dug deeper to see how the models were solving the puzzles. They found that the models that learned to solve the small puzzles (by being trained on them) were actually cheating a little. They relied almost entirely on the shape of the pieces (the bumps and holes) and mostly ignored the picture inside the piece. When the researchers removed the shapes and gave them plain square pieces, these "trained" models fell apart, dropping from 97% accuracy to just 10%. This suggests they hadn't learned to truly understand the image; they had just memorized the shape rules.
In the end, JigShape reveals a hard truth: current AI models are still terrible at geometric reasoning. They can recognize a cat in a photo, but they struggle to figure out how a hundred pieces of that cat fit together in 3D space. While one model showed a glimmer of hope on small tasks, the "scaling cliff" proves that as the complexity increases, these systems hit a wall. The paper suggests that for AI to truly understand the physical world, we need to build new ways for them to combine visual clues with geometric rules, rather than just memorizing patterns. For now, the jigsaw puzzle remains a challenge that even the smartest robots can't quite master.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.