← Latest papers
🤖 AI

AffectAI-Capture: A Reproducible Multimodal Protocol for Small-Group Meeting Research

This paper introduces AffectAI-Capture, a reproducible multimodal protocol designed to collect synchronized data from four-person meetings using diverse sensors like eye tracking and wearables, organized around a unified event timeline to support affective and behavioral research.

Original authors: Meisam Jamshidi Seikavandi, Alice Modica, Anna Obara, Fabricio Batista Narcizo, Tanya Ignatenko, Ted Vucurevich, Jesper Bünsow Boldt, Paolo Burelli, Andrew Burke Dittberner

Published 2026-05-20
📖 5 min read🧠 Deep dive

Original authors: Meisam Jamshidi Seikavandi, Alice Modica, Anna Obara, Fabricio Batista Narcizo, Tanya Ignatenko, Ted Vucurevich, Jesper Bünsow Boldt, Paolo Burelli, Andrew Burke Dittberner

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you want to understand how a group of four friends solves a puzzle together. You could just watch them and listen to them talk, but that only tells you what they said. To truly understand the how and why—who is paying attention, who is stressed, who is leading, and how they feel inside—you need a much more sophisticated toolkit.

This paper introduces AffectAI-Capture, which is essentially a "recipe book" or a strict instruction manual for setting up a high-tech meeting room to study human groups.

Here is a breakdown of what they are doing, using simple analogies:

1. The Goal: Filling the Gap

Think of previous research as having two different types of cameras:

  • Type A: Great for recording meetings (like a security camera), but it only sees the outside behavior (talking, moving).
  • Type B: Great for measuring feelings (like a smartwatch), but it's usually used on one person at a time, not in a group.

AffectAI-Capture is trying to build a bridge. It combines the "meeting room" setup with the "smartwatch" sensors so researchers can see the group dynamic and the individual's internal state at the exact same time.

2. The Setup: A High-Tech Stage

The researchers designed a specific "stage" for four people to sit around a table. They don't just let them chat; they give them specific "scripts" (tasks) to act out, like:

  • The Secret Puzzle: Everyone has different clues, and they must share them to solve a mystery.
  • The Negotiation: They have to bargain over resources.
  • The Brainstorm: They generate ideas and pick the best one.
  • The Trust Game: A mini-game about cooperation and fairness.

These scripts are like standardized "scenes" in a play. By making everyone act out the same scenes, researchers can compare how different groups react to the same pressure.

3. The Sensors: A Swarm of Eyes and Ears

To capture everything, they surround the table with a "swarm" of devices:

  • Smart Glasses: Each person wears glasses that track exactly where their eyes are looking (like a laser pointer on their gaze).
  • Wearable Bands: Each person wears a band that measures their heartbeat, sweat (stress), and breathing.
  • The Camera Rig: Instead of one camera, they use seven cameras. Four are mounted upside down under the table to get a clear view of faces without obstruction, and three are high up to see the whole room.
  • Microphones: They use special microphones that listen to the room and individual voices separately, like having a dedicated microphone for every singer in a choir.

4. The "Golden Clock": Keeping Time

The biggest challenge in this kind of research is synchronization. If the eye-tracking camera is 0.1 seconds slower than the heart-rate monitor, the data gets messy.

The authors treat time like a master conductor in an orchestra. They don't just hope all the devices agree on the time; they build a "safety net" of multiple clocks.

  • They record a "master timeline" (a digital logbook) that marks every event.
  • They use redundant signals (like a visual flash or a specific sound) to check if the cameras and sensors are still in sync.
  • If a device glitches, they have enough "backup time stamps" to fix the data later, rather than throwing the whole experiment away.

5. The "Recipe Book" (Data Organization)

They organize all this messy data using a system called BIDS (Brain Imaging Data Structure). Think of this as a very strict filing cabinet system.

  • Raw Data: The original, unedited video and sensor files are kept safe in one folder (like the original film negatives).
  • Processed Data: The cleaned-up, analyzed versions go in another folder.
  • Metadata: A detailed "instruction manual" explains exactly how the data was collected, so anyone else can repeat the experiment exactly the same way.

6. What They Have Done So Far (and What's Next)

The paper is honest about its current status:

  • The Bench Test: They have tested the equipment in a lab without people. They used robot heads (called Head and Torso Simulators) to check if the microphones work and if the cameras stay in sync.
  • The Result: The audio is clear, and the cameras are synchronized.
  • The Next Step: They haven't run the full experiment with real human participants yet. That is the "ongoing work."

7. Why This Matters

This isn't just about recording meetings. The authors suggest this setup could also be used for Sign Language research. Because sign language uses hand movements in the space around the body, having multiple cameras and eye-tracking (since eye contact is crucial in sign language) makes this protocol very useful for translating or studying signed communication, too.

In short: AffectAI-Capture is a rigorous, reproducible blueprint for building a "super-lab" where researchers can finally study how groups of people think, feel, and interact with a level of detail that was previously impossible.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →