VGGSounder: Audio-Visual Evaluations for Foundation Models
This paper introduces VGGSounder, a comprehensively re-annotated multi-label test set designed to overcome the labeling and alignment limitations of the original VGGSound dataset, thereby enabling more reliable and precise evaluation of audio-visual foundation models through detailed modality analysis and a new confusion metric.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a robot to understand the world by showing it videos. You want the robot to know that if it sees a drum being hit and hears a thump, it should say, "That's a drum!"
For a long time, researchers used a giant library of videos called VGGSound to test these robots. But the authors of this paper, Daniil Zverev and his team, discovered that this library was like a messy, poorly organized attic. It had three big problems that made it impossible to tell if the robots were actually smart or just lucky:
- The "One Label" Trap: The library forced every video to have only one tag. But in real life, a video might show a dog barking while a car honks. The old library would just pick one, ignoring the other. It's like describing a pizza with pepperoni and mushrooms as just "a pizza with cheese."
- The "Confusing Names" Problem: Some labels were twins. One tag might say "dog barking" and another "dog bow-wow." The robots got confused because they were essentially the same thing, but the test treated them as different.
- The "Ghost" Problem: Sometimes the library claimed a sound was in the video when it wasn't actually visible, or vice versa. For example, a video might be labeled "orchestra," but you can only hear the music, not see the instruments. The old test didn't care about this mismatch.
The Solution: VGGSounder
The team built a new, upgraded library called VGGSounder. Think of this as renovating that messy attic into a high-tech, perfectly organized museum.
Here is what they did differently:
- Multi-Labeling: Instead of forcing one tag, they let every video wear as many tags as it needs. A video can be tagged "drums," "guitar," and "singing" all at once.
- Modality Tags: They added a special "checklist" for every single tag. They asked: Is this thing visible? Is this thing audible? This lets them test if a robot is good at seeing, good at hearing, or good at doing both.
- Meta-Labels: They added "warning signs" for tricky situations. If a video has background music, a voice-over, or is just a static photo with sound, they flagged it. This helps researchers see if a robot gets distracted by noise.
The Big Discovery: The "Distraction" Effect
The most surprising thing the team found was how the robots behaved when they were given both eyes and ears.
They invented a new score called "Modality Confusion." Imagine you are good at solving a puzzle using only your eyes. Then, someone hands you a second puzzle piece (the sound) and suddenly you can't solve the first one anymore. That is Modality Confusion.
- The Finding: Many of the newest, most advanced "Foundation Models" (the super-smart AI robots) actually got worse when they were allowed to listen and watch at the same time.
- The Bias: These robots seemed to be "blind" to sound. If you showed them a video of a dog barking but only let them hear it, they failed. But if they could see the dog, they succeeded. When they tried to use both, the sound often confused them, and they ignored it entirely, relying only on what they saw.
Why This Matters
The paper argues that we can't just throw a bunch of data at a robot and hope it learns. We need a better test (VGGSounder) that tells us exactly how the robot is thinking.
- Old Test (VGGSound): Like a driving test where the road is foggy and the signs are missing. You might pass by luck, but you don't know if you're a safe driver.
- New Test (VGGSounder): Like a driving test with clear signs, a co-pilot who points out hazards, and a report card that tells you exactly which senses you used to make your decisions.
The authors conclude that while these new AI models are impressive, they often struggle to combine sight and sound effectively. They tend to ignore the audio if they can see the video, which is a major flaw for machines that are supposed to understand the world holistically. VGGSounder is the tool designed to expose these flaws so engineers can fix them.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.