← Latest papers
💻 computer science

GAViD: A Large-Scale Multimodal Dataset for Context-Aware Group Affect Recognition from Videos

This paper introduces GAViD, a large-scale multimodal dataset of 5,091 videos enriched with contextual metadata and action cues to address the scarcity of data for group affect recognition, alongside the proposed CAGNet model which achieves state-of-the-art performance in recognizing context-aware group emotions.

Original authors: Deepak Kumar, Abhishek Pratap Singh, Puneet Kumar, Xiaobai Li, Balasubramanian Raman

Published 2026-04-20
📖 5 min read🧠 Deep dive

Original authors: Deepak Kumar, Abhishek Pratap Singh, Puneet Kumar, Xiaobai Li, Balasubramanian Raman

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are walking through a busy city park. You see a group of friends laughing, a couple having a tense argument, and a team of strangers cheering for a street performer.

If you just looked at one person's face, you might think, "That guy looks happy." But if you look at the whole group, you realize the "happy" guy is actually the one being teased, and the group's overall mood is actually "playful teasing," not pure joy.

This is the core problem the paper GAViD tries to solve. It's about teaching computers to understand the mood of a whole group, not just the individuals in it.

Here is the story of the paper, broken down into simple parts:

1. The Problem: The Computer is "Blind" to Context

For a long time, computers have been great at reading a single person's face (like spotting a smile). But when it comes to groups, they are often confused.

  • The Missing Puzzle Pieces: Imagine trying to guess the plot of a movie by only looking at a single frame where everyone is standing still. You'd miss the shouting, the laughter, the setting (is it a wedding or a funeral?), and the history of what just happened.
  • The Data Gap: Until now, researchers didn't have a big enough "library" of video clips that included video, sound, and context (what is actually happening) all labeled together. Most existing datasets were like a photo album with no captions.

2. The Solution: The GAViD Dataset (The "Super Library")

The authors built GAViD (Group Affect from ViDeos). Think of this as a massive, high-definition library containing 5,091 video clips of real people interacting in the real world ("in-the-wild").

What makes this library special? It doesn't just show the video; it comes with three layers of information for every clip:

  1. The Visuals: The video itself.
  2. The Audio: What they are saying and the tone of their voices.
  3. The Context (The Secret Sauce): This is the big innovation. They used a smart AI (like a super-advanced librarian) to write a description of the scene.
    • Example: Instead of just seeing "people talking," the context tag says, "A group of friends arguing playfully about a sports game."
    • They also have human annotators (over 100 real people) who watched the clips and labeled the mood (Happy, Sad, Angry, etc.) and the type of interaction (Cooperative, Hostile, Neutral).

The Analogy: If previous datasets were like a silent movie, GAViD is a movie with subtitles, a soundtrack, and a director's commentary explaining exactly what's going on.

3. The Brain: CAGNet (The "Context-Aware Detective")

Having the library is great, but you need a detective to read it. The authors built a new AI model called CAGNet.

  • How it works: Imagine a detective who doesn't just look at the suspect's face.
    • Visual Detective: Looks at facial expressions.
    • Audio Detective: Listens to the tone of voice (is it a shout or a whisper?).
    • Context Detective: Reads the "scene description" (e.g., "This is a funeral").
  • The Magic: CAGNet uses a special "gated fusion" mechanism. Think of this as a traffic light system.
    • If the video is blurry (bad visuals), the traffic light turns green for the Audio and Context detectives to take the lead.
    • If the audio is noisy, the Visual detective takes charge.
    • It dynamically decides which clue is most important at any given moment to figure out the group's mood.

4. The Results: Solving the Mystery

The team tested CAGNet on their new library.

  • The Score: It got about 63% accuracy, which is a very strong score for such a difficult task.
  • Why it matters: When they tested the model without the context (just video and audio), it made mistakes.
    • Example: A group of people shouting. Without context, the AI thinks "Angry." With context ("They are cheering for a goal"), the AI correctly thinks "Happy/Excited."
  • Comparison: Other modern AI models (like generic video chatbots) tried to solve this but scored much lower because they didn't have this specific "context-aware" training.

5. Why Should We Care? (The Real-World Impact)

Why do we need a computer to understand group moods?

  • Better Customer Service: Imagine a smart store that notices a group of friends is getting frustrated and alerts a staff member to help before they leave.
  • Team Building: Analyzing how teams interact in meetings to see if they are collaborating well or if there is hidden tension.
  • Public Safety: Detecting if a crowd is turning from "excited" to "panicked" or "hostile" in real-time.
  • Entertainment: Creating video games or movies that react to the emotional energy of the players or audience.

Summary

The paper is about building a better dictionary for human emotions in groups.

  1. They collected a huge, diverse library of videos (GAViD) that includes sound and AI-generated descriptions of the scene.
  2. They built a smart detective (CAGNet) that learns to weigh the video, the sound, and the story context together.
  3. They proved that when you give an AI the context (the "story"), it becomes much better at understanding what a group of people is actually feeling.

It's a step toward computers that don't just "see" people, but actually "understand" the social dynamics around them.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →