← Latest papers
💬 NLP

Benchmarking and Mechanistic Analysis of Vision-Language Models for Cross-Depiction Assembly Instruction Alignment

This paper introduces IKEA-Bench, a comprehensive benchmark for evaluating Vision-Language Models on cross-depiction assembly instruction alignment, revealing that architectural family outweighs parameter count in performance while identifying disjoint visual subspaces and text-driven reasoning as key mechanistic bottlenecks.

Original authors: Zhuchenyang Liu, Yao Zhang, Yu Xiao

Published 2026-04-03
📖 4 min read☕ Coffee break read

Original authors: Zhuchenyang Liu, Yao Zhang, Yu Xiao

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to build a complicated piece of IKEA furniture. You have the paper instructions (which are just drawings with no words) and a video of someone else building it. Now, imagine you have a super-smart robot assistant that can look at your video and the paper instructions to tell you, "Yes, you're doing it right!" or "Wait, you're using the wrong screw!"

This paper asks a simple question: Can current AI robots actually do this?

The authors built a test called IKEA-Bench to find out. They took 29 different IKEA furniture items, paired the paper diagrams with real video footage, and asked 19 different AI models (the "brains" behind the robots) to solve puzzles about the assembly.

Here is what they found, explained with some everyday analogies:

1. The "Translation Gap" (The Depiction Gap)

The biggest problem is that the AI is terrible at connecting the dots between the drawing and the video.

  • The Analogy: Imagine the paper diagram is a cartoon and the video is a live-action movie. They show the exact same scene, but they look nothing alike. One is a simple black-and-white sketch; the other is a colorful, messy room with real hands and wood.
  • The Result: The AI models are like people who speak two different languages but don't have a dictionary. They can understand the cartoon or the movie, but they can't translate between them. Even the smartest AI models only got about 65% right on the easiest tasks, which isn't good enough to trust with your furniture.

2. The "Text Trap"

The researchers thought, "Maybe if we describe the drawings in words, the AI will understand better?"

  • The Analogy: It's like trying to help a friend find a specific book in a library. You show them a picture of the book (the diagram). They can't find it. So, you write a description: "It's a red book with a dragon on the cover."
  • The Surprise: Instead of helping, the description actually distracted the AI! When the AI read the text description, it stopped looking at the picture of the drawing and started relying only on the words. But the real problem wasn't understanding the words; it was matching the picture of the drawing to the video of the assembly. The text acted like a shiny object that made the AI forget to look at the visual clues it actually needed.

3. The "Brain vs. Body" Problem

The paper dug deep into why the AI fails using a "mechanistic analysis" (basically, looking under the hood of the AI's brain).

  • The Analogy: Think of the AI as having two different "eyes."
    • Eye A looks at the drawings.
    • Eye B looks at the video.
    • The study found that Eye A and Eye B are looking at completely different worlds. They don't even speak the same visual language. The AI's "brain" (the part that makes decisions) is trying to match them, but the inputs are so different that it's like trying to match a fingerprint to a snowflake.
  • The Bottleneck: The AI is actually pretty good at understanding the instructions (the text or the drawing). The real weak link is understanding the video. The AI struggles to figure out what is happening in the messy, real-world video, no matter how smart the rest of the system is.

4. Bigger Isn't Always Better

The researchers tested AI models of all sizes, from small (2 billion parameters) to huge (38 billion parameters).

  • The Analogy: It's like buying a car. You might think a bigger, more expensive car (more parameters) will drive better. But in this case, a newer model of a smaller car drove better than an older, massive truck.
  • The Finding: The type of AI architecture (the "engine design") mattered much more than the sheer size of the model. A newer, smarter design beat an older, bigger one every time.

The Bottom Line

If you want an AI assistant to help you build furniture by watching you and reading the manual:

  1. Current AI isn't ready yet. It gets confused by the difference between a sketch and a real video.
  2. Adding text descriptions doesn't help. It actually makes the AI lazy and stops it from looking at the pictures.
  3. The solution isn't just "bigger brains." We need to teach the AI's "eyes" to see that a sketch and a photo are the same thing, even if they look totally different.

The authors conclude that until we fix this specific "visual translation" problem, our AI assistants will keep telling us we're building the furniture correctly when we're actually putting the legs on upside down!

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →