LLM-Based Selection of Incongruent Verbal and Nonverbal Behavior for Virtual Humans
This paper proposes a taxonomy for verbal-nonverbal mismatches based on Ekman's framework, investigates the use of large language models to select contextually appropriate incongruent behaviors for virtual humans, and validates their effectiveness in producing intended observer effects through a human-subject study.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the quiet spaces between spoken words, human communication is often doing its most complex work. When we talk, we do not just transmit information; we project our feelings, our social standing, and our hidden intentions through our faces, our eyes, and our posture. These silent signals, known as nonverbal behaviors, can reinforce what we say, but they can also contradict it. A person might say "I am fine" while their face tightens with worry, or offer a compliment with a tone that suggests mockery. This ability to let the body betray the mouth, or to mask true feelings behind a polite smile, is a fundamental part of how humans navigate relationships, manage conflict, and signal trust. For decades, researchers have tried to teach computers to mimic this complexity, building virtual agents that can move and gesture like people. However, most of these digital characters have been limited to a simple rule: if the words are happy, the face should be happy. They struggle to understand that sometimes, the most honest thing a person can do is to say one thing while their body says another.
A team of researchers at Northeastern University set out to change this limitation. They wanted to see if they could teach a virtual human to generate these subtle, contradictory behaviors automatically, using a type of artificial intelligence known as a large language model. These models are powerful systems trained on vast amounts of text, capable of understanding context and nuance. The researchers asked a specific question: if they gave the computer a conversation and a description of the situation, could it figure out when the speaker should look unhappy while saying something nice, or look friendly while delivering bad news? They did not just want the computer to guess; they wanted to see if the computer could make these choices in a way that felt real to human observers.
To test this, the team first created a map of the different ways people mismatch their words and their actions. They identified four main reasons why a person might do this. Sometimes it is irony, where a speaker uses praise to insult someone or criticism to show affection. Other times, it is about saving face, where a person softens a harsh critique with a warm smile to avoid hurting feelings, or forces a polite smile to hide their own anger. There is also deception, where a person tries to hide a secret emotion, and emotional regulation, where someone tries to suppress a feeling that is bubbling up inside. Using these categories, the researchers designed eight specific scenarios, such as a junior employee talking to a boss who just increased their workload, or a mother giving feedback to her teenage son.
The researchers then tested three different ways of asking the computer to generate the behavior. In the first method, they gave the computer only the spoken sentence. In the second, they added the history of the conversation leading up to that sentence. In the third and most detailed method, they provided the conversation history plus a written description of the relationship between the speakers and the social setting. The computer, powered by a model called Claude, was asked to decide what the character's face and eyes should do for each sentence. The results were clear: when the computer was given only the words, it mostly made the character look happy when the words were happy and sad when the words were sad. But when the researchers provided the full context—the power dynamics, the history, and the social pressure—the computer began to choose behaviors that contradicted the words. In the scenario where a junior employee had to agree to more work from a demanding boss, the computer selected a smile that was small and restrained, accompanied by a brief look away and a tightening of the lips. It understood that the employee was saying "yes" while feeling "no."
To see if these computer-generated choices actually worked, the researchers turned the digital instructions into video clips of a virtual human and showed them to people. In the first study, they asked participants to watch pairs of videos and decide which speaker had a more positive attitude. When the virtual human said something positive but looked negative, people consistently rated the attitude as less positive than when the human looked happy. When the human said something negative but smiled warmly, people rated the attitude as more positive than when the human looked angry. The second study went deeper, asking people to rate the specific emotions they saw. The results showed that the nonverbal cues the computer selected were powerful enough to shift how people felt about the speaker's emotions. If the computer chose a frown to go with a happy sentence, people reported feeling that the speaker was actually angry or sarcastic. If the computer chose a warm smile to go with a sad sentence, people reported feeling compassion rather than anger.
The study confirms that providing an artificial intelligence with the full social context of a conversation allows it to generate nonverbal behaviors that are far more human-like than those based on words alone. The computer did not just repeat the emotion of the sentence; it interpreted the situation and decided that a mismatch was the most honest response. While the researchers noted that their system still needs work to handle gestures and voice tone, and that the videos were short, the core finding is significant. It suggests that for virtual humans to be truly convincing, especially in roles like training therapists or negotiating, they must be able to let their bodies tell a different story than their mouths. The computer learned that sometimes, the truth is not in the words, but in the silence between them.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.