← Latest papers
🤖 AI

Knowledge-Guided Multimodal Reasoning over Interacting Streams for Video-Level Ambivalence and Hesitancy Recognition

The paper introduces PRISM-AH, a knowledge-guided multimodal framework that recognizes video-level ambivalence and hesitancy by modeling cross-modal dissonance over time and fusing structured evidence from a large language model, achieving a macro F1 of 0.6133 on a public test set.

Original authors: Podakanti Satyajith Chary, Barath Parthiban, Pranesh Velmurugan, Adeeba Khan, Nagarajan Ganapathy

Published 2026-07-29
📖 3 min read☕ Coffee break read

Original authors: Podakanti Satyajith Chary, Barath Parthiban, Pranesh Velmurugan, Adeeba Khan, Nagarajan Ganapathy

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to guess how someone is feeling just by watching a short video of them talking. This isn't just about spotting a smile or a frown; it's about catching those tricky, mixed-up moments where a person says "I'm totally fine!" but their voice cracks, their eyes dart away, and they fidget nervously all at once. Scientists call this state "ambivalence" or "hesitancy." It's that confusing middle ground between "yes" and "no," and it's a huge reason why people often give up on trying to change their health habits, like quitting smoking or starting a diet. The challenge for computers is that these mixed signals are like tiny, fleeting sparks scattered across a person's face, voice, and words. If you just mash all that data together into one big pile, the computer gets confused and misses the subtle drama happening in real-time.

Enter PRISM-AH, a new way of teaching computers to spot these mixed feelings. Instead of just gluing all the video, audio, and text data together and hoping for the best, the researchers built a system that acts more like a detective than a calculator. First, a lightweight "scout" scans the video to find the exact split-seconds where the person's face, voice, and words are disagreeing with each other. It then writes a short, structured report of these moments, highlighting the conflict. Next, a powerful "detective" (a large language model) reads that report. But here's the twist: the detective doesn't just guess; it follows a specific rulebook of "expert clues" provided by human psychologists. By combining the scout's raw data with the detective's expert reasoning, the system gets much better at spotting hesitation.

The results are impressive. On a standard test of 152 hidden videos, this new method achieved a score of 0.6324 (measured by a metric called macro F1). To put that in perspective, the previous best attempt, which just used a standard AI without this special detective logic, only scored 0.2827. That means the new approach is 2.24 times more accurate than the old baseline.

However, the paper is careful to point out what doesn't work. The researchers explicitly ruled out the idea that simply adding more complex features or trying to adapt the AI to specific people (like knowing their age or gender) helps. In fact, they found that conditioning the system on participant attributes actually made it worse when facing new people it hadn't seen before. They also discovered a surprising quirk in how the system handles missing information: while the system uses video, audio, and text, the "text" (what the person is actually saying) is the most critical piece of the puzzle. If you remove the language, the performance drops significantly. But here's the twist: removing the audio does not reduce performance, suggesting that audio is not essential for the decision-making process in this specific setup.

The study suggests that the secret sauce isn't just having more data, but having the right kind of reasoning. The "detective" part of the system only worked when it was given the specific rulebook of expert cues; without that knowledge, the system's performance fell back down to the level of the basic model. The authors conclude that while their current system is a major step forward, the most promising path for future improvements lies in making the language understanding even stronger, since words are the primary driver of the decision. They didn't claim to have solved the problem of reading human emotions forever, but they did prove that a structured, knowledge-guided approach is far superior to just throwing everything into a blender.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →