← Latest papers
💻 computer science

Unveiling Fine-Grained Visual Traces: Evaluating Multimodal Interleaved Reasoning Chains in Multimodal STEM Tasks

This paper introduces StepSTEM, a rigorous graduate-level benchmark with 283 problems and a novel step-level evaluation framework designed to assess fine-grained cross-modal reasoning in multimodal large language models, revealing that current state-of-the-art models still heavily rely on textual shortcuts rather than genuine visual-textual integration.

Original authors: Jing Jin, Hao Liu, Yan Bai, Yihang Lou, Zhenke Wang, Tianrun Yuan, Juntong Chen, Yongkang Zhu, Fanhu Zeng, Xuanyu Zhu, Tao Feng, Yige Xu

Published 2026-05-12
📖 4 min read☕ Coffee break read

Original authors: Jing Jin, Hao Liu, Yan Bai, Yihang Lou, Zhenke Wang, Tianrun Yuan, Juntong Chen, Yongkang Zhu, Fanhu Zeng, Xuanyu Zhu, Tao Feng, Yige Xu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot how to solve a difficult science puzzle, like figuring out how a bridge holds up or calculating the orbit of a satellite. You show the robot a picture of the bridge and a text description of the problem.

The Problem: The "Cheat Code" Robot
Currently, the smartest AI robots (called Multimodal Large Language Models) are like students who are great at reading but terrible at looking. When you give them a science problem with a diagram, they often ignore the picture entirely. They just read the text, guess the answer based on their memory, and say, "I solved it!"

The paper calls this "modality collapse." It's like a student taking a math test who sees a graph but decides to just read the question's words and ignore the graph because it's easier. Existing tests often let them get away with this because they only check if the final answer is right. If the robot guesses the right number without looking at the picture, the test says, "Good job!" But the robot didn't actually learn how to use the visual clues.

The Solution: STEPSTEM (The "Strict Teacher")
The authors created a new, very strict test called STEPSTEM. Think of it as a "graduate-level" exam for robots that forces them to show their work.

  1. The Trap: They designed 283 tricky problems where you cannot solve them without looking at the picture. The text alone is a dead end; the picture holds the key. It's like a treasure hunt where the map (the image) is the only way to find the treasure, and the clues (the text) are useless without it.
  2. The Requirement: The robot isn't just allowed to give an answer; it has to walk through the solution step-by-step, alternating between writing text and drawing pictures. It's like asking a student to solve a physics problem by writing an explanation, then drawing a force diagram, then writing the next step, then drawing a new diagram, and so on.
  3. The Grading: Instead of just checking the final number, the authors built a "smart grader" that looks at every single step.
    • Did the robot look at the right part of the picture? (They use "bounding boxes" like digital sticky notes to check this).
    • Did the text match the drawing?
    • Did the robot follow a logical path, or did it just jump to a conclusion?

The Results: The Robots Are Still Learning
When they ran their top-tier robots through this strict test, the results were surprising:

  • The "Text-Only" Experts: The most powerful robots that can only write text (like Claude Opus 4.6 and Gemini 3.1 Pro) got about 38% of the answers right. They were good at the text parts but completely failed at the drawing parts because they literally cannot draw.
  • The "All-Rounder" Robots: The robots that can both read and draw (like Gemini 2.5 Flash Image) did better at drawing, but they still struggled to get the final answer right. They could draw a picture, but they often got the logic wrong.
  • The Big Gap: The paper found that even the best robots are far from being true "visual thinkers." They rely too much on text and haven't learned to truly combine seeing and thinking.

The Takeaway
The paper argues that we need to stop just asking robots "What's the answer?" and start asking, "Show me how you got there."

Imagine a detective. If a detective solves a crime but didn't look at the fingerprints or the crime scene photos, we wouldn't trust their solution, even if they guessed the right name. STEPSTEM is the tool that forces the AI detective to look at the evidence (the images) and explain how that evidence led to the conclusion, step-by-step.

The authors conclude that while AI is getting smarter, it still has a long way to go before it can truly "see" and "think" about science problems the way a human expert does. They hope this new test will help researchers build robots that don't just guess, but actually understand the visual world.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →