← Latest papers
🤖 AI

MultivationBench: A Benchmark for Multimodal Sequential Motivation Reasoning

The paper introduces MultivationBench, a new benchmark based on psychological frameworks that evaluates the ability of multimodal large language models to perform sequential motivation reasoning in story-driven visual narratives, revealing that current models struggle to maintain consistent reasoning across dynamic contexts despite their static recognition capabilities.

Original authors: Kawai Chung, Chunkit Chan, Yauwai Yim, Yuxuan Liu, Haochen Shi, Weiqi Wang, Qing Zong, Tianshi Zheng, Yixuan Fu, Kai Chung Wong, Hao Liang, Yifan Gao, Xi Yang, Janet Hui-wen Hsiao, Yangqiu Song

Published 2026-07-30
📖 4 min read☕ Coffee break read

Original authors: Kawai Chung, Chunkit Chan, Yauwai Yim, Yuxuan Liu, Haochen Shi, Weiqi Wang, Qing Zong, Tianshi Zheng, Yixuan Fu, Kai Chung Wong, Hao Liang, Yifan Gao, Xi Yang, Janet Hui-wen Hsiao, Yangqiu Song

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are watching a movie, but instead of just seeing what happens, you are trying to guess why the characters are doing it. This is the heart of "social intelligence," a superpower humans have that lets us understand the hidden feelings and goals behind our actions. Scientists call this "motivation reasoning." It's not just about seeing someone pick up a phone; it's about figuring out if they are calling for help, checking the time, or looking at a photo of a lost loved one. Usually, we don't guess this from a single frozen picture. We watch the story unfold, gathering clues over time. If a character looks sad in scene one, then sees a photo in scene two, our guess about their motivation changes. This paper dives into whether the newest, super-smart computer brains (called Multimodal Large Language Models) can do this same kind of "story detective work." While these computers are great at looking at a picture and describing it, or reading a story and answering questions, this research asks a harder question: Can they keep track of a character's changing feelings as a whole story plays out, just like a human does?

The researchers behind this study, led by a team from the Hong Kong University of Science and Technology and Amazon, decided to build a giant test to find out. They created something called MultivationBench. Think of it as a massive library of 1,000 short visual stories, like a mix of movie clips and social media posts, containing over 4,000 specific moments where a character does something. But this isn't just a test of "what happened"; it's a test of "why it happened." To make the test fair and scientific, the team used two famous psychological rulebooks: Maslow's Hierarchy of Needs (which sorts human desires from basic survival like food and safety to higher goals like love and self-improvement) and Reiss's 16 Basic Desires (which looks at specific personal drives like curiosity, honor, or family).

The team asked the computer models to watch these stories frame-by-frame. At every step, the models had to guess the character's motivation based only on what they had seen so far. The tricky part? As the story continued, new information would appear that completely changed the reason for an earlier action. For example, a character might reach for a phone, and at first, it looks like they are just checking a message (a "Cognitive" need). But later, the story reveals they are looking at a picture of their family while feeling lonely at a party, which means the real reason was actually about "Love and Belonging." The computer had to be smart enough to say, "Wait, I was wrong before; now I know the real reason."

The results were a bit of a wake-up call for the tech world. The paper found that while these advanced AI models are decent at guessing motivations in short, simple stories, they really struggle when the story gets longer. Out of the 1,000 stories tested, the models could rarely keep their reasoning consistent from start to finish. In fact, when asked to get every single motivation right in a whole story, the best models only succeeded less than 1% of the time. The study suggests that these AI brains are like students who are great at memorizing facts but terrible at understanding the flow of a plot. They tend to get stuck on the first thing they see and fail to update their thinking when new evidence arrives. They often "hallucinate" reasons that sound plausible but aren't supported by the story, or they miss the deeper, more complex emotional reasons that require connecting dots across the entire narrative.

The researchers also compared the AI to actual humans. The humans, who were graduate students, were significantly better at the task, especially when the stories were long and complex. This gap shows that while AI is getting better at seeing and reading, it still lacks the "social brain" needed to understand how our motivations shift and evolve over time. The paper concludes that we are still far from having computers that can truly understand the messy, changing reasons behind human behavior in a story. It's a reminder that understanding why people do what they do is a lot harder than just knowing what they did.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →