← Latest papers
🤖 AI

M3VerseM^3-Verse: A "Spot the Difference" Challenge for Large Multimodal Models

This paper introduces M3VerseM^3-Verse, a comprehensive benchmark designed to evaluate and improve Large Multimodal Models' ability to reason about dynamic object state changes within consistent spatial contexts through paired video observations, revealing current limitations and proposing an effective baseline for multi-state perception.

Original authors: Kewei Wei, Bocheng Hu, Jie Cao, Xiaohan Chen, Zhengxi Lu, Wubing Xia, Weili Xu, Jiaao Wu, Junchen He, Mingyu Jia, Ciyun Zhao, Ye Sun, Yizhi Li, Zhonghan Zhao, Jian Zhang, Gaoang Wang

Published 2026-05-26
📖 5 min read🧠 Deep dive

Original authors: Kewei Wei, Bocheng Hu, Jie Cao, Xiaohan Chen, Zhengxi Lu, Wubing Xia, Weili Xu, Jiaao Wu, Junchen He, Mingyu Jia, Ciyun Zhao, Ye Sun, Yizhi Li, Zhonghan Zhao, Jian Zhang, Gaoang Wang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: The "Spot the Difference" Game for AI

Imagine you are playing a classic "Spot the Difference" game with two pictures. You look at Picture A, then Picture B, and you have to find what changed. Maybe a cat moved from the sofa to the floor, or a cup was knocked over.

Now, imagine teaching a computer to play this game, but instead of two static pictures, you show it two short videos of a room. In the first video, the room is tidy. In the second video, someone has rearranged the furniture and moved some objects. The computer has to watch both videos and answer questions like, "Where did the red vase go?" or "Did the lamp get turned on?"

This is exactly what the paper M3-Verse is about. The authors created a new, very difficult test to see if modern AI models (called Large Multimodal Models, or LMMs) are actually good at tracking changes over time, or if they are just guessing.

The Problem: AI is Good at "Now," but Bad at "Then vs. Now"

The paper argues that while AI is amazing at looking at a single photo and describing it (like saying, "That's a dog on a rug"), it struggles when asked to compare two different moments in time.

  • Current AI: It's like a tourist who takes a photo of a room, looks away, and then tries to remember what changed when they look back. They often forget the details.
  • Human Intelligence: We naturally compare the past to the present. If a bird flies out of a tree, we instantly notice the difference. The paper says AI is missing this "biological" ability to detect subtle changes between two distinct scenes.

The Solution: A New Test Called M3-Verse

To fix this, the researchers built M3-Verse, a giant test suite with two parts:

  1. The Simulation Lab (M3-Verse-Sim): They used a video game engine (AI2-THOR) to create 270 virtual rooms. They programmed robots to walk around, then they secretly moved objects (like opening a drawer or moving a chair) and recorded the "before" and "after" videos. They generated nearly 3,000 questions about these changes.
    • Analogy: Think of this as a perfectly controlled video game level where the rules are strict, and every detail is known.
  2. The Real World (M3-Verse-Real): They filmed 29 real rooms in actual houses and offices. Humans physically moved objects and filmed the changes. They created 552 questions for these.
    • Analogy: This is the messy, real-life version where lighting changes, cameras shake, and things aren't perfectly aligned.

The test asks the AI to do four things:

  • Spatial Understanding: Where is the object?
  • Temporal Understanding: What happened first, and what happened second?
  • Attribute Recognition: Did the color or shape change?
  • Reasoning: Why did it change, or what does that mean?

The Results: AI is Struggling

The researchers tested 16 of the smartest AI models available (including big names like GPT-5 and Gemini). The results were surprising:

  • Humans: Scored about 90%. We are great at this.
  • Top AI Models: Scored between 30% and 47%. Even the best models are barely better than random guessing on the hardest questions.
  • The Gap: The paper calls this a "stark performance gap." Current AI is not yet smart enough to reliably track how a scene changes from one moment to the next.

The Diagnosis: Why is the AI failing?

The researchers didn't just say "AI failed." They tried to figure out why using a clever trick called the Text-Oracle Probe.

  • The Experiment: They took the videos, had an AI describe every single frame in text (like a detailed transcript), and then asked a text-only AI to answer the questions based only on that text.
  • The Discovery: When the AI didn't have to process the raw video but could just read the text description, its score went way up.
  • The Conclusion: The AI's "eyes" (vision encoder) are actually working fine; it can see the changes. The problem is its "brain" (the reasoning part). When the information is spread out over a long video, the AI forgets the important details. It's like reading a book where the important clues are buried in thousands of pages of boring text; the AI gets lost in the noise and forgets the clue.

The Takeaway

The paper concludes that we need a new generation of AI that doesn't just "see" images but can truly remember and compare them over time.

  • Current State: AI is like a camera that takes great photos but has a terrible memory.
  • Future Goal: We need AI that acts more like a human detective, capable of holding a mental map of a room, watching it change, and spotting the difference between the "before" and "after" states.

The authors released this test (M3-Verse) to the public so other researchers can use it to build better, more "aware" AI systems that can understand our dynamic, changing world.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →