Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding
This paper introduces Auto-AEG, a scalable pipeline that addresses the data scarcity bottleneck in Open-Vocabulary Audio Event Grounding by combining programmatically synthesized clips with multi-model pseudo-labels and reinforcement learning to significantly enhance the temporal localization capabilities of Large Audio-Language Models.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are teaching a super-smart robot to listen to the world. You've already taught it to be a great storyteller; if you play a recording of a storm, it can tell you, "That sounds like heavy rain and thunder." But there's a tricky part: the robot is terrible at knowing exactly when things happen. It might say, "I hear rain," but it can't tell you if the rain started at 2:03 or 2:15. On the other hand, old-school sound detectors are like strict librarians: they can pinpoint the exact second a sound starts and stops, but they only recognize a tiny, pre-approved list of words. If you ask them about a "squeaky door," they might just say, "I don't know that word," even if the sound is right there.
This paper tackles the messy middle ground: teaching a smart robot to not only understand what a sound is but also to point its finger at the exact moment it happens, even if the sound is something it has never heard of before. The big problem is that teaching this skill is incredibly hard because it takes forever to hire humans to listen to hours of audio and mark down every single second a sound starts and stops. It's like trying to teach someone to play the piano by having them watch a million hours of videos without ever letting them touch the keys. The researchers wanted to find a way to teach the robot without needing a million human teachers.
Enter Auto-AEG, a new method that acts like a clever, automated tutor. Instead of waiting for humans to label every sound, the researchers built a pipeline that creates its own practice material. First, they use a computer program to mix and match sounds like a DJ, creating thousands of fake audio clips where the computer knows exactly when every sound starts and stops because it put them there. This gives the robot a "cold start"—a perfect, no-mistakes foundation to learn the basics of timing.
But the robot needs to learn from real life, too, where sounds are messy and overlap. Here's where the magic happens: the researchers let the robot practice on real-world audio (like recordings from the internet) where the timing labels are a bit fuzzy. Instead of just correcting the robot like a teacher, they use a technique called Reinforcement Learning. Think of this like a video game: the robot tries to guess the timing, and if it gets it close enough, it gets a "reward point." If it guesses wildly wrong, it gets no points. Over time, the robot learns to aim for those reward points, figuring out how to be precise even when the training data isn't perfect.
The paper shows that this two-step approach works wonders. When they tested their trained robots on a brand-new, difficult test set called AEGBench (which they also created to make sure the test was fair and covered tricky situations like overlapping sounds or very long noises), the results were impressive. The robots improved their ability to pinpoint sound timing by huge margins—up to 73.9% better than before for the larger model. They didn't just get better at guessing; they learned to be precise.
Crucially, the authors found that simply showing the robot more real-world audio wasn't enough. If they just kept teaching the robot with standard lessons on real audio, it didn't get much better. It was the "reward game" (the reinforcement learning) that turned the messy, imperfect data into a superpower. This suggests that the key to teaching robots to understand time in sound isn't just having more data, but having a smarter way to learn from it. They also proved that this method works on different sizes of robots, from smaller ones to massive ones with 30 billion parameters, and that it helps them stay good at understanding sounds in general, not just at finding them.
In short, this paper suggests that we don't need to wait for humans to label every sound in the universe to build better audio robots. By mixing perfect synthetic practice with a "reward-based" learning game on real-world data, we can teach these models to hear the world with much sharper timing, opening the door for them to understand complex, real-life soundscapes much more effectively.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.