← Latest papers
💻 computer science

EXAM2^2: Extending\underline{Ex}tending Audio\underline{A}udio $Understanding$ $in$ Multilingual\underline{M}ultilingual $and$ Multimodal\underline{M}ultimodal $Analysis$

This paper introduces EXAM2^2, a comprehensive multilingual and multimodal benchmark designed to evaluate and advance audio understanding across six languages and diverse audio-visual scenarios, demonstrating significant performance gaps in current models and the effectiveness of a newly proposed fine-tuned model in addressing these challenges.

Original authors: Jiawen Wang, Xiaoxue Gao, Zi Haur Pang, Nancy F. Chen

Published 2026-08-26
📖 4 min read☕ Coffee break read

Original authors: Jiawen Wang, Xiaoxue Gao, Zi Haur Pang, Nancy F. Chen

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The world of sound is vast and varied, filled with the human voice, the crash of a wave, and the melody of a song. For decades, computers have struggled to make sense of these auditory signals, often treating them as mere noise rather than meaningful information. In recent years, a new generation of artificial intelligence has emerged, capable of listening to audio and answering questions about what it hears. These systems, known as large audio language models, have shown remarkable promise in understanding speech and identifying sounds. However, a significant gap remains in how we test them. Most existing tests focus only on English and rely solely on audio, ignoring the fact that in the real world, sound rarely happens in a vacuum. It is almost always accompanied by a visual scene, and the meaning of a sound can change depending on the language spoken and the context seen.

To bridge this gap, researchers have introduced a new evaluation tool called EXAM2. This benchmark is designed to test how well artificial intelligence can understand audio when it is combined with visual images and spoken in multiple languages. The team behind the project, comprising experts from institutions in Germany, China, Japan, and Singapore, recognized that current tests were too narrow. They built a comprehensive dataset containing thousands of questions that require a computer to listen to an audio clip, look at an image, and then choose the correct answer from a list of options. The questions cover three main types of sound: human speech, environmental noises, and music. Crucially, these questions are not just in English but are translated into five other languages: German, Spanish, Japanese, Malay, and Chinese. This allows the researchers to see if a model's understanding holds up when the language changes or when it must connect what it hears with what it sees.

The researchers constructed this benchmark by carefully selecting audio recordings from various sources, ensuring they represented real-world scenarios rather than synthetic data. For every multiple-choice question, they generated a corresponding image to serve as a visual clue. This process involved a rigorous pipeline of quality control, where human experts reviewed the audio, the text, and the generated images to ensure they were accurate and relevant. The final dataset includes nearly 5,700 questions, over 22,000 images, and more than 135,000 translated instances across the six languages. This massive collection serves as a rigorous training ground and testing field for artificial intelligence, pushing these systems to reason across different types of information simultaneously.

When the researchers tested state-of-the-art artificial intelligence models on this new benchmark, the results revealed both strengths and significant weaknesses. The most capable models, particularly those that could process both audio and images together, performed better than those that relied on sound alone. This suggests that visual context helps the computer disambiguate sounds, making it easier to understand what is happening in a scene. However, the study also found a clear disparity in performance based on language. Models generally performed much better on Western languages like English, German, and Spanish compared to Eastern languages like Japanese and Chinese. This gap was especially noticeable when the task involved identifying environmental sounds or music, indicating that current systems are still heavily biased toward high-resource languages and struggle with the nuances of other linguistic and cultural contexts.

To address these challenges, the team developed their own lightweight model, which they fine-tuned using their new multilingual and multimodal data. This model, trained to handle all six languages and both audio and visual inputs, showed substantial improvements over existing baselines. It achieved a significant boost in accuracy, particularly in the speech and music categories, and demonstrated that training on multiple languages simultaneously helps the model generalize better across different linguistic settings. The findings suggest that while artificial intelligence is becoming better at listening, true understanding requires the ability to see and speak the world's languages. The researchers conclude that for these systems to become truly robust, future development must prioritize diverse languages and the integration of visual and auditory information, moving beyond the limitations of English-only, audio-only testing.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →