Foundation Model Embeddings Meet Blended Emotions: A Multimodal Fusion Approach for the BLEMORE Challenge
This paper presents a 12-encoder multimodal fusion system for the BLEMORE Challenge that achieves 6th place by combining specialized face, audio, and body-language models with the novel application of Gemini Embedding 2.0, demonstrating that frozen prosody layers and personalized salience thresholds are critical for recognizing blended emotions.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are at a crowded party. Someone walks up to you, and they aren't just feeling one thing. They might be 70% excited but 30% nervous. Or maybe they are 50% angry and 50% sad. This mix of feelings is called a "blended emotion."
The BLEMORE Challenge is like a high-stakes game where computers try to guess these mixed feelings just by watching short video clips of people. The tricky part isn't just guessing what emotions are there, but also figuring out how strong each one is (the "salience").
Here is how the team from the University of Zurich built a "super-detective" system to win 6th place in this challenge, explained in simple terms.
1. The Detective Squad (The Multimodal Ensemble)
Instead of hiring just one detective, the team built a squad of six different experts. Each expert looks at the video from a different angle, and then they all sit around a table to vote on the final answer.
- The Face Expert (S4D-ViTMoE): This detective only looks at the person's face. They are trained to spot tiny muscle twitches that show happiness or fear.
- The Body Language Experts (TimeSformer & VideoMAE): Sometimes a face is a poker face, but the body screams the truth. These experts watch the person's shoulders, hands, and posture. If someone is clenching their fists, they know it's anger, even if the face is calm.
- The Sound Expert (Wav2Vec2): This one listens to the voice. But here's the trick: in this challenge, people aren't speaking words; they are making sounds like gasps, laughs, or growls. The team realized that the "middle layers" of this AI (like the middle chapters of a book) are best at understanding the rhythm and pitch of these sounds, while the "top layers" are too busy trying to understand words. So, they ignored the word-hunters and only used the rhythm-hunters.
- The "Super-Brain" (Gemini Embedding 2.0): This is the star of the show. It's a massive, general-purpose AI (like a super-smart robot that has seen the whole internet). The team used it for the first time in this field. Amazingly, this robot only needed to see the first 2 seconds of a video to guess the emotions correctly. It's like a detective who can solve a crime just by glancing at the front door, without needing to search the whole house.
- The Backup Team: They also added a few other pre-trained tools provided by the challenge organizers to make sure they didn't miss anything.
2. The Voting System (Late Fusion)
Once all six experts give their opinion (e.g., "I think it's 60% fear and 40% joy"), the team doesn't just pick the loudest voice. They use a weighted vote.
Think of it like a jury. Some jurors are more experienced than others. The team found that the "Super-Brain" (Gemini) and the "Face Expert" were the most reliable, so they got the most votes. The others got fewer votes, but their input still helped correct mistakes.
3. The "Threshold" Problem (The Shaky Ruler)
This is the most interesting part of their discovery. To turn the AI's fuzzy guesses (like "70% fear") into a final answer, the system uses a "ruler" called a threshold.
- The Problem: The team found that this ruler was unstable. For one actor, the ruler needed to be set very low to catch a subtle emotion. For another actor, the ruler needed to be set very high.
- The Analogy: Imagine trying to guess how loud a person is whispering. If you are talking to a quiet person, a whisper sounds loud. If you are talking to a loud person, the same whisper sounds quiet. The "ruler" changes depending on who is speaking.
- The Result: This "personal style" of each actor was the biggest reason the computer struggled. Even the best AI couldn't perfectly guess the mix because every human expresses emotions differently.
4. The Big Wins
The team learned three major lessons:
- Don't reinvent the wheel: For the sound part, they didn't need to teach the AI everything from scratch. Just picking the right "middle layers" of an existing AI worked better than training a new one.
- Big brains are fast: The giant "Super-Brain" (Gemini) was so smart it could guess emotions correctly from just a 2-second clip, proving that huge AI models are ready for this kind of work.
- Humans are unique: The hardest part of the challenge wasn't the computer's intelligence; it was that humans are all different. A computer can learn the rules, but it can't easily predict how you specifically will express your feelings.
The Final Score
By combining all these experts and voting together, the team achieved a score of 0.279, beating the previous best attempts. They proved that if you want to understand mixed emotions, you need to look at the face, the body, the sound, and use a super-smart AI to tie it all together.
In short: They built a team of specialists who combined their unique skills to read the complex, messy, and very human world of blended emotions.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.