← Latest papers
🤖 AI

InCarEmo: A Multimodal Dataset for In-Cabin Emotion Recognition and Driver State Monitoring

The paper introduces InCarEmo, a comprehensive multimodal dataset integrating RGB, infrared, audio, and dialogue text from scripted in-cabin scenarios to advance driver emotion recognition, fatigue detection, and distraction monitoring while addressing the limitations of existing visual-only datasets.

Original authors: Hao Yang, Yanyan Zhao, Kewei Zhao, Hongbo Zhang, Tian Zheng, Yusheng Liu, Xing Fu, Bichen Wang, Yu Zhang, Hao He, Zhen Wu, Xuda Zhi, Yongbo Huang, Bing Qin

Published 2026-07-17
📖 3 min read☕ Coffee break read

Original authors: Hao Yang, Yanyan Zhao, Kewei Zhao, Hongbo Zhang, Tian Zheng, Yusheng Liu, Xing Fu, Bichen Wang, Yu Zhang, Hao He, Zhen Wu, Xuda Zhi, Yongbo Huang, Bing Qin

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are driving a car, but instead of just a steering wheel and pedals, your vehicle is starting to have a personality. It wants to know if you're happy, tired, or distracted so it can help keep you safe. This is the world of "affective computing"—a fancy term for teaching computers to understand human feelings. Think of it like a super-smart co-pilot that doesn't just read the road, but also reads you. For a long time, scientists have tried to build this co-pilot by looking at your face or listening to your voice. But there's a problem: most of the training data they've used is like a black-and-white movie. It only shows the face or only hears the voice, and it often ignores the fact that when we drive, we are usually talking to someone or something else. It's like trying to understand a joke by only looking at the person's mouth, without hearing the punchline. To make a truly safe and empathetic car, we need to understand the whole picture: the face, the voice, the words, and even the mood of the conversation.

Enter InCarEmo, a new project by researchers from Harbin Institute of Technology and SERES that is like handing the car's computer a full-color, surround-sound movie of real driving life. The team realized that existing datasets were missing a crucial piece of the puzzle: the actual conversation happening inside the car. They built a massive new library of data called InCarEmo, which captures drivers not just looking at the road, but chatting about traffic, planning trips, or discussing the car itself. They recorded these moments using high-definition cameras (both regular and infrared, which sees in the dark), high-quality microphones, and even transcribed the exact words being spoken.

The paper introduces this dataset as a toolkit for three main jobs: figuring out what emotion a driver is feeling (like anger or anxiety), spotting if they are getting too tired to drive, and noticing if they are distracted by their phone or a drink. To test if this new data works, the researchers built a lightweight AI model named CAMEL. They found that when the AI used all the information together—sight, sound, and words—it was much better at guessing the driver's state than when it used just one sense. For example, in low-light conditions where a camera might struggle, the audio and text clues helped the AI stay accurate. The study also created a special English version of the data to help researchers from different countries work together, though they noted that translating emotions across languages is tricky. Ultimately, the paper suggests that by giving AI a richer, more realistic view of the driver's world, we can build cars that are not just smart, but truly understanding and safer for everyone on the road.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →