Multimodal Ambivalence and Hesitancy Recognition via Cross-Attention and Gated Fusion
This paper presents a multimodal framework for recognizing ambivalence and hesitancy in video, which fuses text, audio, and visual features via bidirectional cross-attention and a Gated Multimodal Unit to achieve a Macro F1 score of 0.7394, significantly outperforming both unimodal baselines and zero-shot models in the ABAW11 challenge at ECCV 2026.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to guess what a friend is feeling just by watching them. Sometimes, it's easy: a big smile means happy, a frown means sad. But what about those tricky moments when someone is torn? Maybe they are saying "I'm fine" with a voice that sounds shaky, while their face looks like they are about to cry. This confusing mix of feelings is called ambivalence or hesitancy. It's that internal tug-of-war between wanting to do something and being scared to do it. Scientists who study computers and human behavior (a field called affective computing) are trying to teach machines to spot these subtle, mixed-up signals. They know that looking at just one clue—like only listening to the voice or only watching the face—is often not enough. To really understand a person, a computer needs to act like a super-smart detective, piecing together words, tone, and expressions all at once to solve the mystery of what someone is truly feeling.
In this paper, a team of researchers named Oussama, Yassine, and Larbi decided to build a super-detective computer to solve exactly this puzzle. They entered a competition called ABAW11, where the goal was to spot ambivalence and hesitancy in video recordings. They started by testing three different "senses" on their own: a brain that reads text transcripts, a brain that listens to audio, and a brain that watches facial videos. They found that the text reader was the best solo detective, getting about 66.6% of the answers right (measured by a score called Macro F1). The audio listener was close behind at 62.8%, and the video watcher was a bit weaker at 57.3%. Interestingly, they tried a famous "zero-shot" AI (one that wasn't trained on this specific data) and it only got 28.3% right, proving that you really need to train your AI specifically for this tricky job.
But the real magic happened when they stopped letting the detectives work alone and made them work as a team. The researchers built a special system where the text, audio, and video brains could talk to each other. They used a technique called cross-attention, which is like giving each detective a walkie-talkie to ask the others, "Hey, I see a shaky voice here, does your face see a nervous expression?" and "I see a hesitant word, does your audio hear a pause?" This allowed them to spot contradictions, like when someone says "yes" but sounds unsure. They also added a Gated Multimodal Unit, which acts like a smart manager. If one detective is looking at a blurry face or hearing bad audio, the manager can say, "Ignore that noisy signal for a second, let's focus on the clear text instead."
After training this team with a smart search tool called Optuna to find the perfect settings, the results were impressive. The team of three working together scored a 73.9% on their test, which is a 11.0% improvement over the best single detective. This suggests that by letting the different senses cross-check each other, the computer can catch the subtle, conflicting clues that make up ambivalence much better than any single sense could alone. The researchers are confident that this method of explicit teamwork between text, sound, and video is the key to unlocking these difficult human emotions, and they have made their code public so others can try it out too.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.