AIR-BENCH Live: An Evolving Safety Benchmark for Foundation Models
This paper introduces AIR-BENCH Live, a self-evolving safety benchmark for foundation models that utilizes an automated pipeline to continuously integrate new government regulations and generate multilingual attack prompts, thereby addressing the limitations of static benchmarks in a rapidly changing AI landscape.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the world of artificial intelligence as a massive, ever-expanding library where robots are learning to write stories, solve problems, and even chat with us. But just like in any library, there are rules about what can be written: no instructions on how to build bombs, no hate speech, and no tricks to steal secrets. Scientists create "safety tests" to see if these robots follow the rules. Think of these tests like a series of riddles or traps designed to trick the robot into breaking the rules. If the robot refuses the trick, it gets a high score; if it falls for it, it gets a low score. The problem is that the world changes fast. New laws are passed, new ways to trick robots are invented, and old riddles become too easy to solve. A safety test from last year might be useless today because the robot has learned to ignore the old tricks, or because a new law has banned something the test never even thought of. This is why researchers are always looking for a way to make these tests "live" and update themselves, so they stay relevant in a world that never stops moving.
This paper introduces a new, self-updating safety test called AIR-BENCH Live. Think of it as a security guard that doesn't just stand at the door with a fixed list of banned items, but instead has a robot assistant that constantly scans the news, reads new laws from around the world, and instantly creates new, tricky riddles to test the AI. The researchers built a pipeline that automatically "scrapes" (reads) government regulations from places like the US, China, and the EU. When it finds a new rule—like a law against stealing a robot's brain or using fake videos to mess with elections—it adds a new category to its safety checklist. Then, a team of AI agents works together to write realistic, multi-language prompts (the riddles) based on these new rules. They use "personas," which are like acting roles, to make the riddles sound like they are coming from real people in different situations, rather than a robot reading a script.
The team tested 14 different AI models using this new, evolving benchmark. They found a huge gap in safety: some models were incredibly safe, refusing almost every trick (scoring 1.00), while others were quite easy to trick (scoring as low as 0.17). Interestingly, the new, modernized prompts were harder than the old ones, causing even the most compliant models to slip up a bit more. The researchers also discovered that most models were slightly less safe when asked in languages other than English, suggesting that keeping robots polite and safe across all human languages is still a work in progress. The paper concludes that while we can't stop the world from changing, we can build safety tests that evolve right alongside it, ensuring that as AI gets smarter, our ability to test its safety gets smarter, too.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.