Assessing AI-generated music detection in real-world broadcast monitoring
This paper introduces BAMM, a real-world dataset of 40 hours of television recordings, to demonstrate that current CNN-based AI-generated music detectors suffer from critical performance degradation and insufficient reliability when applied to real broadcast monitoring conditions compared to clean or synthetic environments.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a music detective trying to solve a mystery in a crowded, noisy room. In this room, there are two types of singers: real humans and a new kind of robot that can sing perfectly by learning from millions of songs. Your job is to spot the robot singer. In a quiet studio, this is easy; the robot's voice has a tiny, invisible "glitch" that only a detective with super-hearing can catch. But what happens when that robot tries to sing while a loud crowd is shouting, a siren is wailing, and the sound is being beamed through a cheap, crackly radio? This is the world of broadcast monitoring. It's the science of listening to what's actually playing on TV and radio, where music is rarely alone. It's often short, hidden behind dialogue, or muffled by poor signal quality. As AI music becomes more common, companies and governments need to know: Can our current "detectives" still find the robot singers when the room is chaotic, or do they get confused and miss them entirely?
This paper, titled "Assessing AI-Generated Music Detection in Real-World Broadcast Monitoring," is a reality check for those detectives. The authors, a team from a music technology lab and a licensing company, realized that most previous tests were like practicing for a storm while standing in a calm, air-conditioned room. They built a new tool called BAMM (Broadcast AI-Music Monitoring), which is a massive library of 40 hours of real television recordings. This isn't a computer simulation; it's actual TV clips where AI music and human music are playing in the background of news, ads, and shows, just as they would in the real world.
The researchers put two different types of "detective" computers to the test. The first one, CNN Clean, was trained only on perfect, high-quality music in a quiet studio. The second, CNN Broadcast, was trained on messy, mixed-up audio that sounded like TV. They tested both on three levels of difficulty:
- The Quiet Studio: Perfect, clear music.
- The Fake TV: Music mixed with speech in a controlled computer simulation.
- The Real TV: The actual 40 hours of real broadcast clips from the BAMM dataset.
Here is what they found. In the quiet studio, both detectives were nearly perfect, catching the robot singers almost 100% of the time. But as soon as they moved to the "Fake TV" scenario, the detective trained only on quiet music (CNN Clean) completely lost its way, dropping its success rate to a mere 34%. The detective trained on messy audio (CNN Broadcast) did better, reaching 66%, proving that training on messy data helps. However, when they moved to the Real TV scenario—the actual, unscripted chaos of real television broadcasts—both detectives stumbled badly. The clean-trained one dropped to 18%, and even the broadcast-trained one only reached 47%.
The most troubling discovery was that in the real TV world, the scores for human music and AI music started to look exactly the same. The detectives couldn't tell them apart anymore. The paper suggests that the "glitches" the computers were looking for in the quiet studio get completely washed out or hidden when the music is short, mixed with talking, or compressed into a low-quality signal. While training on messy data makes the detectors a bit tougher, it isn't enough to solve the problem. The authors conclude that our current methods are not yet ready to reliably spot AI music in the real world of television and radio, and we need much smarter detectives to handle the noise.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.