← Latest papers
🤖 AI

StemBind: When MLLMs Get Lost Between Rules and Instances in Abstract Visual Reasoning

The paper introduces StemBind, a diagnostic benchmark that reveals a persistent "binding gap" in multimodal large language models where, despite correctly identifying visual rules, they frequently fail to map those rules to specific instances during abstract visual reasoning, a limitation that neither model scaling nor explicit thinking modes currently resolve.

Original authors: Xixiang He, Baiqi Wu, Xingming Li, Ao Cheng, Qiyao Sun, Xuanyu Ji, Qingyong Hu

Published 2026-06-02
📖 4 min read☕ Coffee break read

Original authors: Xixiang He, Baiqi Wu, Xingming Li, Ao Cheng, Qiyao Sun, Xuanyu Ji, Qingyong Hu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are taking a very tricky visual puzzle test, like the kind used to measure human intelligence (think Raven's Progressive Matrices). You look at a grid of shapes, figure out the hidden rule connecting them, and pick the right piece to finish the pattern.

Now, imagine a super-smart AI taking this test. Here is the weird thing the paper STEMBIND discovered: The AI often knows the rule perfectly, but still picks the wrong answer.

It's like a student who can recite the entire textbook definition of a math formula and explain exactly how it works, but then fails to plug the numbers into the formula correctly to get the final answer.

Here is a breakdown of what the researchers found, using simple analogies:

1. The Problem: The "Lost in Translation" Gap

Current AI tests usually just look at the final answer: "Did you get it right or wrong?" This is like a teacher grading a math test by only looking at the final number, ignoring whether the student actually understood the steps.

The researchers built a new test called STEMBIND to look inside the AI's brain. They asked the AI three questions about the exact same puzzle:

  1. Perception: "What shapes do you see?" (Can you see the ingredients?)
  2. Rule: "What is the pattern?" (Do you know the recipe?)
  3. Full: "Which piece completes the puzzle?" (Can you cook the dish?)

The Shocking Discovery:
On 22 out of 24 different AI models tested, the AI could correctly describe the shapes and correctly name the rule. But when asked to pick the final answer, it failed.

  • Analogy: Imagine a chef who can perfectly describe a cake recipe and list every ingredient, but when asked to bake it, they put the cake in the oven upside down. They know the theory, but they can't bind that theory to the specific action of baking.

2. The Specific Culprit: The "Mapping" Step

The researchers broke the thinking process down into four steps (based on a famous psychology theory):

  • S1 (Encode): Seeing the shapes.
  • S2 (Infer): Figuring out the rule.
  • S3 (Map): <-- The Problem Spot Connecting the rule to the specific answer choices.
  • S4 (Apply): Picking the final answer.

They found that the AI gets stuck at S3. It's like having a map and a destination, but forgetting which street leads to which building. The AI knows the rule is "rotate 90 degrees," but when looking at the four options, it can't figure out which one is the result of that rotation.

3. What Doesn't Fix It?

The paper tested two common ways people think AI gets better:

  • Making the AI Bigger (Scaling): They tried using much larger, more powerful versions of the AI.
    • Result: The bigger AI got better at seeing the shapes and naming the rules, but it still failed to pick the right final answer. It's like giving a student a bigger library; they know more facts, but they still can't solve the specific problem.
  • Making the AI "Think" (Thinking Mode): They let the AI write out its thoughts before answering (like showing its work).
    • Result: This actually made things worse. The AI got better at describing the picture, but it got worse at picking the rule and the final answer.
    • Analogy: It's like a student who starts talking to themselves while solving a problem. They describe the picture beautifully, but the extra talking confuses them so much they forget the rule and pick the wrong answer.

4. Why This Matters

The paper argues that we need to stop just ranking AI models by their final score. We need to diagnose where they break down.

The main takeaway is that today's top AI models are great at seeing and understanding rules, but they are terrible at applying those rules to specific instances. They get "lost" between knowing the rule and choosing the answer.

In short: The AI isn't blind, and it isn't stupid about the logic. It just has a glitch in the "translation" step where it tries to turn a known rule into a specific choice. Fixing that "binding" gap is the next big challenge for AI researchers.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →