A Production-Oriented Framework for Evaluation of SFX Generation
This paper introduces a production-oriented evaluation framework for reference-guided sound effects generation that addresses the limitations of existing text-to-audio assessments by defining nine industrial requirements and a two-stage protocol to systematically compare heterogeneous methods across metrics like reference alignment, diversity, and perceptual identity preservation.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a sound designer for a video game. You have one perfect recording of a crow cawing, but your game needs that same crow to sound slightly different every time it appears, without losing its "crow-ness." You need a digital magic wand that can take your original sound and spin off a hundred new versions: some louder, some with a different echo, some that sound like they are in a cave, but all still sounding like that specific crow.
For a long time, the tools to do this were like a box of mismatched keys. Some tools were great at making any sound from a text description (like "crow"), but they couldn't remember your original recording. Others were great at copying a sound exactly, but they couldn't change it. The researchers in this paper decided to build a new "test drive" track to see which tool actually works best for this specific job.
The Big Test: A Shared Race
The authors set up a fair race using a standard set of 50 sound categories (like "crow," "rain," or "dog") from a public library called ESC-50. They took five different high-tech audio generators and asked them all the same question: "Here is a reference sound. Make me 10 new versions of it."
They didn't just listen to see if the sounds were "good." They measured three specific things:
- Identity: Did the new sounds still sound like the original reference? (If you asked for a crow, did it sound like a crow, or did it drift into a chicken?)
- Diversity: Were the 10 new versions actually different from each other, or were they just carbon copies?
- Realism: Did the sounds feel natural, or did they have that weird, robotic "synthetic" texture?
The Winner: The "Balanced" Contender
After running the numbers, one model called AudioX emerged as the strongest all-rounder for this specific job.
Think of the other models as specialists who are great at one thing but terrible at another.
- AudioLDM was the "wild card." It produced the most different-sounding variations (a diversity score of 0.43), but it often forgot what the original sound was supposed to be. Its "identity alignment" score was only 0.39, meaning the new sounds sometimes drifted too far from the original.
- T-Foley tried to control the timing and energy of the sound, but it struggled to keep the sound realistic. It had the worst "Fréchet Audio Distance" (a measure of how weird the sound is) at 24.53, and humans rated its identity preservation very low at 1.89 out of 5.
- A2SB was the "surgeon." It didn't try to rewrite the whole song; it just fixed small, masked holes in the audio. When it did this, it was amazing, scoring a near-perfect 4.81 out of 5 for identity. But the paper argues this isn't a fair comparison for generating whole new sounds, because it only changed tiny bits while leaving the rest of the original file untouched.
AudioX, however, hit the sweet spot. It kept the original sound's identity very strong (an alignment score of 0.59) while still creating enough variety (a diversity score of 0.27). Humans rated its identity preservation at 3.37 out of 5, which was the highest among the models that tried to generate the whole sound from scratch. It didn't go wild like AudioLDM, but it didn't get stuck in a loop either.
The Trade-Off: You Can't Have It All
The paper suggests a clear trade-off: if you want sounds that are wildly different from the original, you might lose the original's "soul." If you want to keep the soul perfectly intact, you might get fewer variations.
The researchers found that AudioX is the best compromise for a production workflow where you need to keep the "character" of the sound while still making it fresh. But they also point out that if you need to do something very specific, like only changing the volume of a specific part of a sound (targeted editing), a different tool called ThinkSound might be better. ThinkSound is good at following instructions like "make this part quieter," but it's not the best at generating a whole new sound from scratch.
What This Means for the Future
The paper doesn't claim that the problem is "solved." Instead, it suggests that we need to stop comparing these tools using generic tests that don't match real-world needs. Just because a tool is good at making music from text doesn't mean it's good at editing a specific sound effect.
The authors argue that for sound designers, the goal isn't just to make "plausible" audio. It's to make audio that respects the original recording, offers useful variations, and fits into a real workflow. Based on their measurements, AudioX currently offers the strongest overall balance for this, but the "best" tool depends entirely on what specific job you are trying to do. The paper concludes that we need a new way to evaluate these tools—one that looks at the whole picture of identity, diversity, and control, rather than just a single score.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.