← Latest papers
💻 computer science

Exploring Audio Hallucination in Egocentric Video Understanding

This paper introduces a systematic evaluation framework and a curated dataset to reveal that state-of-the-art audio-visual language models suffer from significant audio hallucinations in egocentric video understanding, often inferring sounds from visual cues rather than actual audio, thereby highlighting the critical need for robust reliability assessments in multimodal systems.

Original authors: Ashish Seth, Xinhao Mei, Changsheng Zhao, Varun Nagaraja, Ernie Chang, Gregory P. Meyer, Gael Le Lan, Yunyang Xiong, Vikas Chandra, Yangyang Shi, Dinesh Manocha, Zhipeng Cai

Published 2026-04-28
📖 4 min read☕ Coffee break read

Original authors: Ashish Seth, Xinhao Mei, Changsheng Zhao, Varun Nagaraja, Ernie Chang, Gregory P. Meyer, Gael Le Lan, Yunyang Xiong, Vikas Chandra, Yangyang Shi, Dinesh Manocha, Zhipeng Cai

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are wearing a camera on your chest, recording your day from your own point of view. This is called an "egocentric" video. Sometimes, the camera shakes, gets blocked by your hand, or points at a blank wall. In these moments, your eyes might be confused, but your ears are still working hard.

This paper is about teaching computers to listen to these videos, but the researchers discovered a funny and dangerous problem: The computers are "hearing" things that aren't there.

Here is the breakdown of their findings using simple analogies:

1. The Problem: The Computer is "Reading the Room" Instead of Listening

Think of a state-of-the-art AI model (a super-smart computer brain) as a detective who is trying to solve a mystery based on a video.

  • The Ideal Detective: Listens to the audio tape and says, "I hear a dog barking."
  • The Flawed Detective (The AI in this paper): Looks at the video, sees a picture of a dog, and immediately says, "I hear a dog barking," even if the audio tape is completely silent or just has wind noise.

The researchers call this "Audio Hallucination." It's like a person who, upon seeing a picture of a lemon, insists they can smell the citrus, even though they are in a room that smells like nothing. The AI is so obsessed with what it sees that it invents sounds to match the picture.

2. The Experiment: The "Sound Quiz"

To prove this, the researchers built a special test, like a pop quiz for these AI detectives.

  • The Setup: They took 300 real-life videos of people doing things (like cooking or walking) and chopped them into short 10-second clips.
  • The Questions: They asked the AI 1,000 specific questions.
    • Type A (The Truth): "What sound is happening right now?" (e.g., "Is the blender running?")
    • Type B (The Trap): "Where is the hissing sound coming from?" (When there is absolutely no hissing sound in the video).

They wanted to see if the AI would say, "There is no hissing," or if it would lie and say, "Oh, that's probably the gas stove," just because it saw a stove in the video.

3. The Results: The AI Failed the Trap

The results were quite poor. Even the smartest AI models (like Qwen2.5 Omni) were terrible at this.

  • When asking about real sounds: The AI got about 56% to 63% of the answers right. It was okay at describing what was actually happening.
  • When asking about fake sounds (The Trap): The AI's accuracy plummeted to between 27% and 39%.

The Metaphor: Imagine a student taking a test. If you ask, "What color is the sky?" they get it right. But if you ask, "What color is the invisible elephant?" they confidently say, "It's blue," because they are guessing based on what they think an elephant should look like, rather than admitting they can't see one.

The AI is essentially "filling in the blanks" with guesses based on the visuals, rather than admitting, "I don't hear that."

4. Two Types of "Fake Hearing"

The researchers sorted these mistakes into two categories:

  1. The "Main Character" Mistake (Foreground): The AI hears a sound that the person in the video should be making.
    • Example: The video shows someone using a blender. The AI says, "I hear a mechanical whirring sound." Even if the audio is broken or silent, the AI assumes the sound is there because the blender is moving.
  2. The "Background Noise" Mistake (Background): The AI hears sounds from the environment that aren't there.
    • Example: The video shows a quiet kitchen. The AI says, "I hear birds chirping outside," just because it sees a window.

5. Why This Matters

The paper concludes that right now, these AI models are unreliable listeners. They rely too heavily on their eyes and not enough on their ears.

If you were to use this technology in a real-world scenario (like helping a robot understand a noisy workshop or a blind person's surroundings), the robot might think it hears a warning siren when it's actually just seeing a red light. The researchers say we need to build better "ear-checking" systems before we can trust these models to understand the world through sound.

In short: The paper shows that today's smartest video-AIs are like people who are so good at reading lips that they forget to listen to the voice, often making up sounds that fit the picture they see.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →