← Latest papers
💻 computer science

PolyComp: A Polycube-based Benchmark for Compositional 3D Spatial Reasoning in Multimodal Models

The paper introduces PolyComp, a procedurally generated benchmark of 120 problems designed to evaluate compositional 3D spatial reasoning in multimodal models, revealing that while advanced models like GPT-5.6 Sol achieve moderate accuracy (50%), others struggle near random guessing levels despite varying costs and presentation formats.

Original authors: Siddharth Patel

Published 2026-08-18
📖 5 min read🧠 Deep dive

Original authors: Siddharth Patel

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The human mind possesses a quiet, almost instinctive talent for taking apart complex shapes and imagining how they fit back together. We can look at a puzzle piece, a broken toy, or a stack of boxes and instantly visualize how the parts relate to the whole, even when we cannot see every surface. This ability, known as spatial reasoning, is a cornerstone of how we navigate the physical world. For decades, scientists have wondered if the new generation of artificial intelligence, which can see images and read text, has developed this same internal sense of space. While these machines excel at recognizing cats in photos or summarizing articles, it remains unclear whether they can truly understand the three-dimensional geometry of objects, or if they are simply guessing based on patterns they have seen before.

A new study introduces a rigorous test designed to settle this question, moving beyond simple picture recognition to probe the core of how machines think about shape and assembly. The researchers created a benchmark called PolyComp, a collection of 120 visual puzzles built from blocky, cube-based structures. In each puzzle, an artificial intelligence is shown a target object made of several connected cubes, viewed from two different angles. It is then presented with four possible pairs of smaller pieces and asked to identify which pair, when rotated and moved, would perfectly combine to recreate the original target. The task requires the model to mentally build a 3D representation of the pieces, simulate how they might turn in space, and verify if they fit together without gaps or overlaps. This is not a test of memory or language, but a pure test of geometric logic.

The study put three of the most advanced artificial intelligence systems to the test: GPT-5.6 Sol, Claude Fable 5, and Gemini 3.1 Pro Preview. Each system was asked to solve the same 120 puzzles, but the puzzles were presented in three different ways to see if the format of the image helped or hindered the reasoning. Sometimes the target and the options appeared together in a single image; other times, the target was shown in one image and the four options were split across four separate images, either with generic labels or descriptive names. The goal was to see if breaking the visual information into smaller chunks made the task easier, or if the models struggled regardless of how the information was displayed.

The results revealed a significant gap between human-level intuition and current machine capabilities. The best-performing model, GPT-5.6 Sol, managed to solve only half of the puzzles correctly, achieving an accuracy of 50 percent. The second model, Claude Fable 5, solved about 39 percent, while the third, Gemini 3.1 Pro Preview, performed near the level of random guessing, getting just 27.5 percent of the answers right. Since there were four options for every puzzle, a computer guessing blindly would be expected to get 25 percent correct. This means the most advanced models were barely better than chance, and the least advanced were essentially guessing. The study also found that the type of shape mattered more than the way the images were presented. The models struggled most with puzzles involving rectangular loops, where the pieces formed a ring-like structure, but performed slightly better on puzzles that looked like simple blocks being split apart or joined together.

Interestingly, the way the puzzles were shown did not make a dramatic difference in the outcome. Splitting the images into separate files or adding descriptive labels did not significantly boost the scores for any of the models. This suggests that the difficulty lies not in the models' ability to see the pictures clearly, but in their inability to construct a stable, internal 3D model of the objects and simulate how they move. The researchers noted that when the models did succeed, it was often because the pieces had a very obvious overall shape or a large, flat surface that made them easy to match. When the puzzles required tracking small hidden corners or figuring out how pieces fit inside a hollow space, the models consistently failed.

The study also highlighted the cost of this kind of deep reasoning. To solve the 360 total questions (120 puzzles presented three ways each), the most capable model consumed a vast amount of computing resources, costing nearly one dollar per question when using standard pricing. The less capable models were cheaper to run but also less accurate. This trade-off suggests that even when we pay for the most powerful systems available, they still lack the fundamental spatial intelligence that humans develop naturally. The researchers released all 120 puzzles and the code used to generate them, inviting others to test future models against this same standard.

Ultimately, this work serves as a clear diagnostic tool, showing that while artificial intelligence has made tremendous strides in understanding language and recognizing objects, it has not yet mastered the ability to mentally manipulate three-dimensional space. The models can describe a cube or identify a shape, but they cannot yet reliably imagine how two separate pieces of a shape would fit together if turned and rotated in the air. Until machines can build these internal 3D maps and simulate transformations within them, their spatial reasoning will remain fragile, prone to errors that a human child would likely solve with ease. The study does not claim that this ability is impossible for machines to learn, but it firmly establishes that current systems have not yet achieved it, leaving a clear path for future research to bridge the gap.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →