← Latest papers
⚡ electrical engineering

AVSD-Scenes: A Dataset for Audio-Visual Description of Urban Scenes

This paper introduces AVSD-Scenes, a novel dataset of 12,291 paired audio-visual descriptions for urban environments generated by combining modality-specific outputs from advanced AI models, which demonstrates superior performance in semantic alignment, cross-modal retrieval, and scene classification compared to unimodal approaches.

Original authors: Dhanunjaya Varma Devalraju, Arshdeep Singh, Mark D. Plumbley

Published 2026-10-02
📖 5 min read🧠 Deep dive

Original authors: Dhanunjaya Varma Devalraju, Arshdeep Singh, Mark D. Plumbley

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Cities are never silent, nor are they ever still. They are a constant, shifting tapestry of sound and sight, where the rumble of a bus, the chatter of a crowd, and the visual sweep of a park bench all happen at once. For decades, scientists have tried to teach computers to understand these environments, but they have often relied on simple labels, like tagging a recording simply as "park" or "street." These labels are useful, but they are thin; they tell us the location but miss the story. They do not capture the specific rustle of leaves, the color of a passerby's coat, or the way a sound echoes off a building. To truly understand a city, a computer needs more than a name; it needs a description that weaves together what is seen and what is heard into a single, coherent narrative.

This is the challenge that researchers at IIT Madras and King's College London have tackled with a new project called AVSD-Scenes. The team set out to create a massive collection of urban scenes, each paired with a rich, natural language description that explains exactly what is happening in both audio and video. They started with an existing library of 12,291 recordings from ten European cities, where every clip contained synchronized sound and video. However, the original library only offered basic category tags. The researchers wanted to fill in the gaps, transforming those simple tags into detailed paragraphs that describe the specific events, objects, and actions occurring in each scene.

To achieve this, they built an automated pipeline that acts like a team of specialized observers. First, they used powerful artificial intelligence models to look at the sound and the video separately. One model listened to the audio, identifying specific sounds like bird chirps or traffic noise, and wrote a short summary of what it heard. Another model watched the video, spotting people, vehicles, and movements, and wrote a summary of what it saw. These two separate accounts were then fed into a third, more advanced language model. This final model acted as an editor, combining the audio notes and the visual notes into a single, unified story. It was instructed to ensure the story made sense, to avoid making things up, and to keep the description grounded in the actual evidence provided by the sound and the image. The result was a dataset where every urban scene is accompanied by a paragraph that reads like a human observation, capturing the full context of the moment.

The researchers then tested whether these computer-generated descriptions were actually useful. They asked a series of questions to see if the descriptions could help a computer understand the world better. One test involved seeing if a computer could use a text description to find the correct video or audio clip from a large library. The results showed that the new multimodal descriptions were significantly better at this task than descriptions based on just sound or just sight. When the computer used the combined description, it could find the matching audio clip much more often, suggesting that the text had successfully learned to capture the essence of the sound.

They also tested if these descriptions could help a computer identify the type of city scene it was looking at, such as distinguishing a busy metro station from a quiet public square. Using only the text descriptions, the system achieved an accuracy of 94.5 percent in identifying the correct scene. When the researchers combined the text with the original audio and video data, the accuracy rose slightly to 95.4 percent. This was a notable improvement over using audio or video alone, which had struggled to reach similar levels of precision. The findings suggest that the written descriptions captured deep, meaningful details about the environment that the raw data alone had missed.

Perhaps the most revealing part of the study was a test designed to see if the computer was just memorizing the scene names or actually understanding the content. The researchers removed the scene labels from the instructions given to the AI during the creation process, forcing it to rely solely on the audio and visual content to generate the descriptions. Even without being told the name of the place, the descriptions remained highly effective at distinguishing between different types of scenes. This indicated that the AI was not simply repeating the label "park" because it was told to; it was genuinely describing the grassy areas, the wooden fences, and the specific sounds of the environment. The descriptions were capturing the reality of the scene, not just the category.

The team also evaluated the quality of the writing itself. They used both automated tools and human judges to rate the descriptions on factors like fluency, grammar, and how well the text matched the audio and video. The human judges found the descriptions to be generally accurate and informative, with high scores for how well they described what was seen and how smoothly they were written. While there were occasional moments where the descriptions could be improved, the overall quality was strong enough to be useful for complex tasks.

This work represents a significant step forward in how machines perceive the world. By moving beyond simple labels to rich, descriptive narratives, the researchers have shown that computers can learn to understand the complex interplay of sound and sight in our daily lives. The dataset, which is now available for others to use, provides a new foundation for building systems that can truly "see" and "hear" the world as we do. It suggests that the future of artificial intelligence in urban environments lies not in better labels, but in better stories.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →