← Latest papers
💻 computer science

GIQ: Benchmarking 3D Geometric Reasoning of Vision Foundation Models with Simulated and Real Polyhedra

This paper introduces GIQ, a comprehensive benchmark utilizing diverse synthetic and real polyhedra to evaluate vision and vision-language foundation models, revealing significant deficiencies in their geometric reasoning capabilities across tasks like 3D reconstruction, symmetry detection, and shape classification.

Original authors: Mateusz Michalkiewicz, Anekha Sokhal, Tadeusz Michalkiewicz, Piotr Pawlikowski, Mahsa Baktashmotlagh, Varun Jampani, Guha Balakrishnan

Published 2026-02-06
📖 5 min read🧠 Deep dive

Original authors: Mateusz Michalkiewicz, Anekha Sokhal, Tadeusz Michalkiewicz, Piotr Pawlikowski, Mahsa Baktashmotlagh, Varun Jampani, Guha Balakrishnan

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a group of very smart robots that have been shown millions of pictures of the world. They are great at recognizing cats, cars, and even writing poems. But the authors of this paper wanted to ask a simple question: Do these robots actually understand the 3D world, or are they just memorizing patterns?

To find out, they created a test called GIQ (Geometric IQ Test). Think of GIQ as a "driver's license exam" for robots, but instead of driving a car, they have to understand shapes like cubes, pyramids, and complex star-shaped objects.

Here is how the paper breaks down, using some everyday analogies:

1. The Test Subjects: The "Polyhedra"

The researchers didn't use random objects like apples or chairs. They used polyhedra—geometric shapes made of flat faces, like dice, soccer balls, or complex star crystals.

  • The Range: They started with simple shapes (like a perfect cube) and moved to incredibly complex ones (like a "Miller's Monster," a shape with 124 faces that looks like a tangled mess of stars).
  • The Setup: They tested the robots in two ways:
    • The Video Game World: Perfect, computer-generated images of these shapes.
    • The Real World: They built actual paper models of these shapes, took them outside, and photographed them in snow, sunlight, and messy living rooms. This is like testing if a robot can recognize a toy in a video game and a real toy on a dusty kitchen table.

2. The Four Challenges

The paper put the robots through four specific tests to see how "geometrically smart" they really are.

A. The "One-Shot" Reconstruction (Can they build it from a photo?)

The Task: Show the robot a single picture of a 3D shape and ask it to build the full 3D model.
The Result: Total Failure. Even the best robots, trained on millions of 3D objects, couldn't get it right.

  • The Analogy: Imagine showing someone a photo of a perfect cube and asking them to build it. Instead of making a cube with straight edges and sharp corners, the robot builds a wobbly, lopsided blob. It couldn't even get the basic "right angles" of a cube correct. When the shapes got more complex, the robots' "builds" fell apart completely.

B. The Symmetry Spotter (Can they see the pattern?)

The Task: Show the robot a shape and ask, "Does this have a center point where it looks the same if you flip it?" or "Does it have a 4-fold rotation (like a cross)?"
The Result: Surprisingly Good.

  • The Analogy: While the robots were terrible at building the shapes, they were actually quite good at spotting the symmetry. It's like a person who can't draw a perfect circle but can instantly tell you if a drawing is symmetrical. The researchers found that the robots' "eyes" (their internal brain layers) naturally picked up on these 3D patterns, even without being explicitly taught to do so.

C. The Mental Rotation Test (Can they imagine turning it?)

The Task: Show the robot two pictures of the same shape, but one is rotated. Ask: "Are these the same object?"
The Result: Struggling.

  • The Analogy: This is like holding a Rubik's cube in your hand, turning it in your mind, and seeing if it matches another cube. The robots got confused easily. When the shapes were simple, they did okay. But when the shapes were tricky or the photo was taken in the real world (with shadows and weird angles), the robots started guessing randomly.
  • The Human Comparison: The researchers asked real humans to take the test. While the best robot matched the average human score, 68% of the actual humans did better than the best robot. The robots just couldn't handle the subtle differences.

D. The Name Game (Zero-Shot Classification)

The Task: Show the robot a picture of a complex shape and ask, "What is the name of this shape?" (e.g., "Is this a Johnson solid or a Catalan solid?").
The Result: Confused and Hallucinating.

  • The Analogy: Imagine showing a robot a complex star-shaped toy and asking, "What is this?" The robot might say, "It's a star!" but then make up details.
    • It might confuse a concave shape (one that caves in) with a convex one (one that sticks out).
    • It might mix up two different shapes stuck together (a compound) with a single complex shape (a stellation).
    • Even the smartest AI assistants (like the ones that power ChatGPT or Gemini) got less than 20% of the complex shapes right. They often "hallucinated" features that weren't there, like inventing a hexagon face on a shape that only had triangles.

3. The Big Takeaway

The paper concludes that while these AI models are amazing at recognizing textures (like "this looks like a cat") or patterns, they lack true geometric understanding.

  • The "Cheat" vs. The "Skill": The robots seem to be "cheating" by memorizing what things look like from their training data, rather than understanding the math of how the shapes are built.
  • The Gap: There is a huge gap between how well these models do on standard tests and how they handle real geometric reasoning. They are like a student who memorized the answers to a math test but doesn't actually know how to do the math.

In short: If you ask these robots to build a house based on a blueprint, they might build a wobbly mess. But if you ask them to point out a symmetrical window, they might get it right. The paper argues that until we fix this "geometric blindness," these robots aren't truly ready to navigate or understand the complex 3D world we live in.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →