← Latest papers
💻 computer science

NARRATE: A Multimodal Real-World Australian Driving Dataset for Human-Centred Explanations in Automated Driving

The paper introduces NARRATE, a multimodal real-world Australian driving dataset featuring 2,050 events annotated with synchronized sensor data and driver-generated explanations to advance human-centred, domain-aware models for automated driving.

Original authors: Ashkan Yousefi Zadeh, Zishuo Zhu, Xiaomeng Li, Andry Rakotonirainy, Sebastien Glaser, Ronald Schroeter, Patricia Delhomme, Zahra Mehraban

Published 2026-08-18
📖 4 min read☕ Coffee break read

Original authors: Ashkan Yousefi Zadeh, Zishuo Zhu, Xiaomeng Li, Andry Rakotonirainy, Sebastien Glaser, Ronald Schroeter, Patricia Delhomme, Zahra Mehraban

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine sitting in the back seat of a car that drives itself. The vehicle slows down, turns, or stops, and you are left wondering why. For the technology to feel safe and trustworthy, it must be able to tell you what it sees and why it is making a move, using words that a human passenger can actually understand. This need for clear communication is the heart of a new field of research focused on making artificial intelligence in cars explainable. Scientists are trying to teach machines not just to drive, but to narrate their decisions in a way that matches how human drivers think. The challenge lies in understanding the gap between what a computer calculates and what a person actually notices and anticipates on the road. To bridge this gap, researchers must first understand how real humans explain their own driving choices in the middle of real traffic.

A team of researchers in Australia has taken a significant step toward solving this puzzle by creating a new collection of real-world driving data called NARRATE. Instead of relying on computer simulations or asking people to guess what they might do in a virtual world, the team recorded thirty-five experienced drivers and driving instructors as they drove normal routes through the streets of Brisbane. The drivers were equipped with a vehicle filled with cameras, lasers, and sensors that captured every detail of the road, the car's movement, and the surrounding environment. As they drove, the researchers asked them to explain their actions in real time, and then again after the trip while watching a video replay of what had just happened. This dual approach captured both the quick, instinctive thoughts drivers have while behind the wheel and the more detailed reflections they offer once the pressure of the moment has passed.

The resulting dataset contains two thousand and fifty specific driving events, each paired with the driver's own words explaining what they saw, what they understood about the situation, and what they thought might happen next. The researchers organized these explanations using a framework that breaks down human awareness into three parts: noticing a detail, understanding what that detail means for the current drive, and projecting what could happen in the future. For example, a driver might notice a traffic light turning yellow, understand that they need to slow down, and project that a car behind them might not stop in time. By tagging the drivers' spoken words with these layers of awareness, the team created a rich map of how human reasoning connects to the physical world.

When the researchers tested whether computer models could learn from this data, they found that machines were quite good at identifying the basic structure of human awareness. If a model read a driver's explanation, it could reliably tell whether the driver was simply noticing something, understanding its meaning, or anticipating a future risk. However, the study also revealed that the task is much harder when it comes to the finer details. While computers could easily recognize broad categories like "following a car" or "stopping at a light," they struggled to distinguish between thirty-two specific types of driving situations based on text alone. Similarly, when asked to generate new explanations from scratch, the computer models produced text that was grammatically correct but often lacked the specific, context-rich nuance that human drivers naturally provide.

The findings suggest that while artificial intelligence can learn the skeleton of human reasoning, it still lacks the deep, intuitive grasp of context that comes from years of driving experience. The dataset shows that human explanations are not just descriptions of actions; they are woven with observations of the environment, interpretations of other road users, and predictions of future events. The researchers concluded that for automated vehicles to truly communicate with their passengers, they need to move beyond simple sensor data and learn to reason in a way that mirrors the human mind. This new dataset provides a crucial foundation for that work, offering a rare glimpse into the real-world thoughts and words of the drivers who know the roads best.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →