← Latest papers
🤖 AI

SeePhys Pro: Diagnosing Modality Transfer and Blind-Training Effects in Multimodal RLVR for Physics Reasoning

The paper introduces SeePhys Pro, a benchmark revealing that current multimodal models struggle with modality transfer from text to images, and demonstrates that apparent gains from multimodal RLVR training often stem from residual textual cues rather than genuine visual reasoning.

Original authors: Kun Xiang, Terry Jingchen Zhang, Zirong Liu, Bokai Zhou, Yueling Tang, Junjie Yu, Jiacong Lu, Shangrui Huang, Heng Li, Likui Zhang, Kunkun Liu, Changzheng Zhang, Yangle Fang, Boqiang Guo, Hui-Ling Zhe
Published 2026-05-12
📖 5 min read🧠 Deep dive

Original authors: Kun Xiang, Terry Jingchen Zhang, Zirong Liu, Bokai Zhou, Yueling Tang, Junjie Yu, Jiacong Lu, Shangrui Huang, Heng Li, Likui Zhang, Kunkun Liu, Changzheng Zhang, Yangle Fang, Boqiang Guo, Hui-Ling Zhen, Dandan Tu, Yinya Huang, Xiaodan Liang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a student how to solve a physics problem. You have two ways to give them the information:

  1. The Textbook Method: You write out the problem in words, describing the setup, the numbers, and the rules.
  2. The Diagram Method: You draw a picture of the setup, label the parts, and write the numbers directly on the drawing.

Intuitively, you'd expect a smart student to solve the problem equally well whether you give them the words or the picture, as long as the information is the same. But this paper, SEEPHYS PRO, discovered that for current AI models, this isn't true. In fact, they get significantly worse when you switch from words to pictures.

Here is a breakdown of what the researchers found, using simple analogies.

1. The "Same Story, Different Format" Test

The researchers created a special test called SEEPHYS PRO. Think of it like a "translation challenge" for AI.

For every single physics problem, they created four versions that tell the exact same story but change how the information is presented:

  • Level 1 (Text Only): "A block weighs 5kg and is on a ramp." (All info is in words).
  • Level 2 (Structure in Image): They draw the ramp and the block, but the text still says "5kg." (The shape is a picture, the numbers are words).
  • Level 3 (Variables in Image): They draw the ramp, the block, and write "5kg" right on the block in the picture. (The shape and numbers are now in the image).
  • Level 4 (Full Picture): The entire problem, including the instructions and numbers, is handwritten on a piece of paper and scanned as one image.

The Discovery:
When the AI models tried these tests, their performance dropped like a stone as the information moved from text to images.

  • The "Reading the Label" Bottleneck: The biggest drop in performance happened at Level 3. This is when the AI had to look at the picture and "ground" the variables (e.g., "Okay, that squiggly line is a wire, and the number '12' next to it means 12 Volts").
  • The Analogy: Imagine a chef who can perfectly follow a written recipe (Level 1). If you hand them a photo of the ingredients with labels (Level 3), they suddenly forget how to cook. They can see the picture, but they can't connect the label "2 cups of flour" to the actual bag of flour in the photo. They get confused by the visual binding.

2. The "Blindfolded Training" Experiment

The researchers then asked a second question: If we train these AIs using Reinforcement Learning (RL) on thousands of physics problems with pictures, will they get better at reading the pictures?

To test this, they set up a "Blind Training" experiment.

  • Normal Training: The AI sees the problem text and the picture.
  • Blind Training: The AI sees the problem text, but the picture is completely blacked out (masked). It's like training a student to solve physics problems while wearing a blindfold, but the teacher keeps giving them the same "correct answer" rewards.

The Shocking Result:
Even though the AI was trained with no pictures at all (just black screens), it still got better at solving the unmasked tests where pictures were present.

  • The Analogy: Imagine a student who practices for a driving test while sitting in a car with the windows painted black. They never see the road. Yet, when they finally take the real test with windows open, they drive better.
  • Why? The researchers found the AI wasn't learning to "see" the road. Instead, it was learning shortcuts. It memorized patterns in the text, the style of the questions, or the typical answers. It learned to guess the right answer based on the "vibe" of the text, not by actually analyzing the visual diagram.

3. The "Magic Trick" of Accuracy

The paper concludes that when we see an AI's score go up, we shouldn't just cheer. We need to ask: Did they actually learn to see, or did they just get better at guessing based on text clues?

  • The "Blind Gain": The AI improved its score even when the training images were useless (black). This proves that the improvement came from the text and the patterns of the dataset, not from visual understanding.
  • The Gap Remains: Even after training, the AI still struggled more with the picture-heavy versions (Level 3 and 4) than the text-only versions. The "modality gap" (the difference between text performance and image performance) didn't close.

Summary

  • The Problem: Current AI models are "text-heavy." They are great at reading physics problems but terrible at "seeing" them, especially when they have to read numbers and labels directly off a diagram.
  • The Trap: Standard training methods (Reinforcement Learning) can make the AI look smarter by improving its text-based guessing skills, even if it never actually learns to interpret the visual evidence.
  • The Lesson: To truly make AI "see," we can't just measure if it gets the right answer. We have to test if it gets the right answer specifically because it looked at the picture, and not just because it recognized the text pattern.

The paper essentially says: Don't be fooled by high scores. If an AI can solve a problem without looking at the picture, it hasn't truly learned to see.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →