← Latest papers
🤖 machine learning

Attention-Guided Dual-Stream Learning for Group Engagement Recognition: Fusing Transformer-Encoded Motion Dynamics with Scene Context via Adaptive Gating

This paper proposes DualEngage, a novel dual-stream framework that fuses transformer-encoded individual motion dynamics with scene-level spatiotemporal context via adaptive gating to achieve high-accuracy group engagement recognition in classroom videos.

Original authors: Saniah Kayenat Chowdhury, Muhammad E. H. Chowdhury

Published 2026-04-14
📖 4 min read☕ Coffee break read

Original authors: Saniah Kayenat Chowdhury, Muhammad E. H. Chowdhury

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a busy classroom as a bustling orchestra. Some students are playing their instruments with passion (highly engaged), some are just tapping their feet occasionally (medium engagement), and others are daydreaming or looking out the window (low engagement).

For a long time, teachers and computers have tried to figure out who is paying attention. Most systems only looked at one musician at a time, checking if their eyes were open or if they were nodding. But in a group setting, the magic isn't just about the individual; it's about the vibe of the whole room. Sometimes a whole class is silent and still, but everyone is deeply focused. Other times, they are moving around, but not really learning.

This paper introduces a new system called DualEngage (Dual = Two). It's like hiring two different detectives to solve the mystery of "Who is engaged?" and then having them compare notes before making a final decision.

The Two Detectives (The Two Streams)

Detective #1: The "Motion Tracker" (The Primary Stream)

  • What they do: This detective zooms in on every single student. They don't just look at faces; they watch how people move. They use a special tool (called Optical Flow) that acts like a "wind sensor" for pixels. It sees the subtle shifts in posture, the fidgeting, the leaning forward, or the turning of the head.
  • The Superpower: It uses a "Transformer" (think of it as a super-smart memory bank) to remember how a student moved over time. Did they lean in at the start of the lecture and stay there? Or did they start fidgeting after five minutes?
  • The Problem: Sometimes, a student is perfectly still because they are too focused. If this detective only looked at movement, they might think, "Oh, no movement means no engagement!" which would be a mistake.

Detective #2: The "Room Vibe" Expert (The Secondary Stream)

  • What they do: This detective steps back and looks at the entire classroom as one big picture. They don't care about individual students; they care about the "group energy."
  • The Superpower: They use a 3D camera brain (a 3D ResNet) to see patterns like: Is the whole class leaning forward together? Are they all looking at the teacher? Is the room synchronized?
  • The Problem: This detective might miss the quiet student in the back who is actually the most engaged, because they are too focused on the big picture.

The Smart Manager: The "Gated Fusion"

Here is the genius part. In the past, computers would just take the notes from both detectives and mash them together equally. But that's like asking a weather forecaster and a traffic cop to decide if you should wear a coat, without letting them talk to each other first.

DualEngage has a Smart Manager (called Softmax-Gated Fusion).

  • How it works: The manager looks at the current situation and asks: "Which detective is more reliable right now?"
    • Scenario A: The room is chaotic, and students are moving a lot. The manager says, "Okay, Detective #1 (Motion) is doing great work. Let's listen to them more."
    • Scenario B: The room is silent and still. The manager says, "Detective #1 is confused because everyone is still. But Detective #2 (Room Vibe) sees that everyone is staring at the board in unison. Let's trust the Room Vibe more!"

The manager dynamically adjusts the "volume" of each detective's voice based on what's happening in the video.

The Results: A Standing Ovation

The researchers tested this system on real classroom videos with groups of students.

  • The Score: It got it right 96% of the time.
  • Why it matters: Previous systems struggled when students were quiet or when the camera angle was tricky. This system figured out that "stillness" can mean "focus" if the whole group is synchronized, and "movement" can mean "distraction" if the group is out of sync.

The Bottom Line

Think of DualEngage as a teacher who has a superpower: they can watch every single student's body language and feel the energy of the whole room at the same time. They know when to focus on the individual and when to focus on the group.

By combining individual motion (how you move) with group context (how the room feels) and using a smart switch to decide which one matters more at any given second, this system can finally understand the complex, messy, and beautiful reality of a real classroom. It's a big step toward helping teachers know exactly when their students are truly learning.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →