← Latest papers
💬 NLP

Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm

This paper proposes "Thinking with Video" as a novel multimodal reasoning paradigm that leverages video generation models like Sora-2 as a unified medium to overcome the limitations of text and image-based approaches, demonstrating through the VideoThinkBench that such models achieve state-of-the-art performance on both vision-centric and text-centric reasoning tasks.

Original authors: Jingqi Tong, Yurong Mou, Hangcheng Li, Mingzhe Li, Yongzhuo Yang, Ming Zhang, Qiguang Chen, Tianyi Liang, Xiaomeng Hu, Yining Zheng, Xinchi Chen, Jun Zhao, Xuanjing Huang, Xipeng Qiu

Published 2026-04-08
📖 6 min read🧠 Deep dive

Original authors: Jingqi Tong, Yurong Mou, Hangcheng Li, Mingzhe Li, Yongzhuo Yang, Ming Zhang, Qiguang Chen, Tianyi Liang, Xiaomeng Hu, Yining Zheng, Xinchi Chen, Jun Zhao, Xuanjing Huang, Xipeng Qiu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

🎬 The Big Idea: From Static Photos to a Moving Movie

Imagine you are trying to solve a tricky puzzle.

  • Old Way ("Thinking with Text"): You write down your thoughts on a piece of paper. It's great, but it's just words.
  • Newer Way ("Thinking with Images"): You draw a picture to help you think. It's better, but the picture is frozen in time. It can't show you how something moves or changes.
  • The New Way ("Thinking with Video"): You don't just draw a picture; you make a short movie of your thinking process. You can draw a line, watch it bounce, see a shape change, and hear the answer spoken out loud, all in one continuous flow.

This paper introduces a new way for AI to think called "Thinking with Video." Instead of just reading or looking at a static image, the AI (specifically a model called Sora-2) solves problems by generating a video that shows its reasoning step-by-step.


🧩 The Problem with the Old Tools

The authors argue that current AI has two main blind spots:

  1. The "Snapshot" Problem: Images are like photos. They capture a single moment. If you want to understand how a ball bounces or how a light beam reflects off a mirror, a single photo isn't enough. You need to see the movement.
  2. The "Language vs. Vision" Wall: Usually, AI treats text (words) and vision (pictures) as two separate languages. It's like having a translator who is great at French and great at Spanish, but struggles to mix them together naturally.

The Solution: Video is the perfect bridge. It contains both moving images and (in this new model) spoken or written text. It's like a multitasking chef who can chop vegetables (visuals) while explaining the recipe (text) at the same time.


🏆 The Test: "VideoThinkBench"

To see if this "Thinking with Video" idea actually works, the researchers built a giant playground called VideoThinkBench. Think of it as a video game level designed to test the AI's brain.

The test had two main types of levels:

1. The Visual Gym (Vision-Centric Tasks)

These tasks require the AI to "see" and "draw" to solve problems.

  • The "Eyeballing" Game: Imagine a geometry test where you have to guess where two lines will cross. The AI doesn't just guess; it draws the lines on a digital whiteboard in the video, extends them, and marks the spot.
    • Result: Sora-2 was amazing here. It beat top competitors by actually drawing the solution, showing it could "visualize" physics and geometry better than models that just stare at a picture.
  • The Maze: The AI had to draw a path through a maze.
    • Result: It was okay at square mazes but struggled with weird shapes (like circles), showing it still has room to grow.

2. The School Exam (Text-Centric Tasks)

These are classic math and logic problems (like "If James buys a plane for $150k...").

  • The Twist: The AI had to solve the math problem by writing the steps on a whiteboard in the video and speaking the final answer.
    • Result: This was the surprise! Sora-2 got about 92% accuracy on hard math problems (MATH dataset). It was almost as good as the best text-only AI models.
    • The Catch: While the answer was usually right, the handwriting in the video was often messy or hard to read. It was like a student who knew the math perfectly but had terrible penmanship.

🔍 How Does It Work? (The Secret Sauce)

The researchers dug deep to figure out why Sora-2 was so smart. They found three interesting things:

  1. It's a "Copycat" Learner (Few-Shot Learning):
    If you show Sora-2 a few examples of how to solve a puzzle before asking it to solve a new one, it gets much better. It's like showing a student three solved math problems before giving them a test. The more examples you give, the smarter it gets.

  2. The "Try Again" Effect (Self-Consistency):
    If you ask Sora-2 to solve the same problem five times and pick the answer that appears most often, it gets much more accurate. It's like asking a group of people a question and taking a vote; the "wisdom of the crowd" helps filter out mistakes.

  3. The "Translator" Secret (Prompt Rewriter):
    This is the most fascinating discovery. When Sora-2 solves a hard text math problem, it seems to have a hidden "translator" inside it. Before the video generation starts, this internal translator rewrites the complex math question into a simple, step-by-step visual script (e.g., "Write '6 + 6 = 12' on the board").

    • Analogy: It's like a movie director who reads a complex script, rewrites it into simple action instructions for the actors, and then the actors perform it. The video generator is the actor; the "rewriter" is the director doing the heavy thinking.

🚀 Why Does This Matter?

This paper suggests a huge shift in how we build AI.

  • Unified Brain: Instead of having one AI for reading, another for seeing, and another for drawing, "Thinking with Video" suggests we can have one AI that does it all. It understands the world by simulating it in motion.
  • Human-Like Thinking: Humans don't just think in static words; we imagine movies in our heads. We visualize a ball bouncing, a car turning, or a math equation being written out. "Thinking with Video" makes AI think more like a human does.

📝 The Bottom Line

The paper shows that video generation models (like Sora-2) are not just fancy video makers; they are powerful reasoning engines. By letting the AI "think" by making a video, it can solve visual puzzles better than ever and surprisingly well at math too.

It's like upgrading from a still camera to a live-action movie studio. The AI isn't just looking at the world anymore; it's acting out the solution.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →