← Latest papers
💻 computer science

CALM-AH: An ABAW11-Calibrated Multimodal Ensemble with Reliability-Gated Multi-Expert Consensus for Video-Level Ambivalence and Hesitancy Recognition

The paper presents CALM-AH, a multimodal ensemble system enhanced by a Reliability-Gated Multi-Expert Consensus (RG-MEC) mechanism that achieves a Macro-F1 score of 0.7771 on the ABAW11 challenge by combining diverse feature branches with a unanimity-gated correction strategy to recognize video-level ambivalence and hesitancy.

Original authors: Wenzhuo Sun, Mingjian Liang, Richard Attfield, Zongyuan Ge, Xuelian Cheng, Pamela Carreno-Medrano

Published 2026-08-03
📖 5 min read🧠 Deep dive

Original authors: Wenzhuo Sun, Mingjian Liang, Richard Attfield, Zongyuan Ge, Xuelian Cheng, Pamela Carreno-Medrano

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are sitting in a room, trying to figure out if someone is truly unsure about a big decision. They might say, "I think I'll go," but their voice wavers, they pause for a long time, and they keep glancing at the floor. Or maybe they say, "I'm definitely sure," while fidgeting nervously. This tricky mix of words, sounds, and body language is called ambivalence (having mixed feelings) or hesitancy (being unsure). It's a subtle human state that happens all the time, from job interviews to therapy sessions, but it's incredibly hard for computers to spot. Unlike a simple smile or a frown, ambivalence is like a secret code hidden in the gaps between words, the rhythm of a voice, and the tiny twitch of a muscle. If we could teach machines to read these clues, we could build better tools to help people make tough choices or understand each other better. This is the challenge that a group of researchers at Monash University set out to solve.

Enter CALM-AH, a new digital detective team designed to crack the case of human uncertainty. The researchers realized that looking at just one clue—like only reading the transcript or only watching the face—is like trying to solve a mystery with a single blurry photo. You need the whole picture. So, they built a system that acts like a super-organized jury. Instead of one judge, they have 15 different "jurors," each looking at a different combination of clues: some listen to the voice, some read the words, some watch the face, and some even track the weird statistical patterns of how the person is behaving.

Here's how the team works: First, they gather all the evidence. They use smart AI tools to turn spoken words into text, analyze the rhythm and pauses in the voice, and scan the video for facial expressions. Then, they mix and match these clues in 15 different ways. For each mix, they pick the best "detective" (a type of computer algorithm) and teach it exactly when to shout "Guilty!" (Ambivalent) or "Not Guilty!" (Not Ambivalent). These 15 detectives then vote on the final answer. If most of them agree, that's the verdict. This part of the system, called CALM-AH, was already pretty good, scoring a 0.7525 on a test designed to see if it works on people it has never met before.

But the researchers didn't stop there. They knew that even smart detectives make mistakes. Sometimes a detective might get confused by a loud noise or a tricky sentence. So, they added a special "Safety Net" called RG-MEC (Reliability-Gated Multi-Expert Consensus). Think of this as a very strict rule for changing the verdict. The system starts with a "Default Judge" who makes the first call. To change that judge's mind, you don't just need one other expert to disagree; you need three specific experts to all agree on the exact same new answer at the same time.

These three experts are:

  1. CALM-AH (the 15-detective team we just met).
  2. AffectGPT (an AI that is really good at understanding emotions and feelings).
  3. A GPT-based Verifier (an AI that is great at understanding the meaning and logic of what is being said).

If the Default Judge says "Not Ambivalent," but CALM-AH, AffectGPT, and the Verifier all shout "Ambivalent!" together, then the system changes its mind. If even one of them is unsure or disagrees, the system sticks with the original judge's decision. This prevents the system from getting confused by a single noisy expert. It's like a club where you can only change the rules if three specific members sign the petition together; one person complaining isn't enough to flip the switch.

The result? This "Safety Net" made the system even sharper. By using this strict, unanimous-vote rule, the final score jumped to 0.7771. The paper shows that this method is much better at spotting real uncertainty without getting tricked by random pauses or accidental noises. The researchers found that simply adding more experts doesn't help; it's about how you organize them. By giving the "Default Judge" a stable anchor and only letting the "Safety Net" override it when everyone agrees, the system becomes more reliable.

The team tested this on a dataset of real interview videos where the people in the test videos were completely different from the people the system was trained on. This is crucial because it proves the system isn't just memorizing faces; it's actually learning to spot the feeling of uncertainty. The paper explicitly rules out the idea that you can just average all the experts' opinions or let a single strong expert overrule the rest. Instead, they found that a strict "unanimity gate" works best. While the system is a big step forward, the authors are careful to say this is a specific solution for this particular challenge, not a magic bullet for every human emotion. But for figuring out when someone is truly on the fence, CALM-AH with its RG-MEC safety net is a very smart new tool in the box.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →