← Latest papers
🤖 AI

The Gordian Knot for VLMs: Diagrammatic Knot Reasoning as a Hard Benchmark

KnotBench introduces a rigorous benchmark using nearly 860,000 knot diagrams to demonstrate that current vision-language models, despite improved performance with "thinking" modes, fundamentally struggle to translate visual knot structures into actionable symbolic reasoning, often failing to outperform random baselines on tasks requiring diagrammatic manipulation.

Original authors: Hao Liu, Jicheng Liu

Published 2026-05-12
📖 5 min read🧠 Deep dive

Original authors: Hao Liu, Jicheng Liu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a very smart robot that can look at a picture of a tangled string and describe it perfectly. It can say, "I see a red loop crossing over a blue loop, which then goes under a green one." It has excellent eyesight.

But here is the catch: if you ask that same robot to untangle the string, or to tell you if two different pictures of knots are actually the same knot just drawn differently, it completely fails. It can describe the knot, but it cannot "play" with it in its mind.

This paper, titled "The Gordian Knot for VLMs," introduces a new test called KnotBench to prove exactly how big this gap is between "seeing" and "doing" for modern AI models.

The Setup: A Digital Knot Garden

The researchers created a massive garden of 858,000 pictures of knots.

  • The Source: They started with 1,951 basic "prototype" knots (the simplest building blocks of knot theory).
  • The Magic: They used a mathematical rule (called a Reidemeister move) to twist, turn, and redraw these knots thousands of times.
  • The Result: You might have a picture of a simple "trefoil" knot with 3 crossings, and another picture of the exact same knot with 20 crossings, looking completely different. Or, you might have two pictures that look almost identical but are actually different knots (like a left-handed glove vs. a right-handed glove).

The computer knows the "true identity" of every single knot because it generated them using math, not human guesswork. This makes the test perfectly fair.

The Test: 14 Ways to Get Stuck

The researchers asked four of the smartest AI models available (including Claude Opus 4.7 and GPT-5) to perform 14 different tasks. They split the tasks into two groups to see where the AI breaks down:

  1. The "Image" Group: The AI looks at a picture of the knot.
  2. The "Text" Group: The AI looks at a list of numbers and letters (a code) that describes the knot.

Here is what they asked the AI to do, using simple analogies:

  • The "Same or Different?" Test (Family A): "Are these two pictures the same knot?"

    • The Trap: Sometimes the pictures look totally different but are the same knot. Sometimes they look almost identical but are different.
    • The Result: When given the text code, the AI was great at this. When given the picture, it guessed randomly, like a monkey throwing darts.
  • The "What Happened?" Test (Family B): "I showed you a knot, then I made one small move (like adding a loop). Which move did I make?"

    • The Trap: This requires the AI to mentally simulate the move. "If I pull this string here, what does the picture look like?"
    • The Result: When given the text code, the AI could figure it out. When given the picture, it failed miserably. One model (GPT-5) just started guessing the most common answer ("R3") over and over again, like a student who memorized the answer key but didn't learn the math.
  • The "Counting" Test (Family C): "How many times do the strings cross?"

    • The Result: The AI could count 8 crossings easily. But once the knot got complex (17+ crossings), the AI's count became a complete mess. It couldn't even count the items it was looking at.
  • The "Translator" Test (Family D): "Here is a picture. Write down the secret code for it."

    • The Result: Zero. None of the models could do this. They couldn't translate the visual picture into the mathematical code, even with their "thinking" mode turned on.

The Big Discovery: The "Perception-Operation Gap"

The paper calls the failure point the "Perception-Operation Gap."

Think of it like this:

  • Perception is looking at a map and saying, "I see a river and a bridge."
  • Operation is looking at that same map and saying, "If I walk across the bridge, I will get to the other side."

The AI models are amazing at Perception. They can describe the knot. But they are terrible at Operation. They cannot take what they see and mentally manipulate it to solve a problem.

Does "Thinking" Help?

The researchers turned on the AI's "thinking mode" (where the model talks to itself before answering).

  • Did it help? Yes, but only a little.
  • Where did it help? Only when the AI was given the text code. If the AI had to look at a picture first, "thinking" didn't fix the problem. It's like giving a calculator to someone who can't read the numbers on the screen; the calculator is useless.

The Conclusion

The paper concludes that current AI models are like photographers who can't be mechanics. They can take a perfect photo of a knot and describe every detail, but they lack the internal "simulator" that humans use to mentally twist and turn objects to understand how they work.

They can see the knot, but they cannot act on it. Until AI builds this internal "mental simulator," it will remain blind to the logic of diagrams, no matter how many pictures it sees.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →