Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs
This paper demonstrates that an autoresearch pipeline powered by Claude Code can autonomously discover novel white-box adversarial attack algorithms that significantly outperform existing methods, achieving up to 100% attack success rates against state-of-the-art LLM safeguards.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a very smart, very stubborn robot guard (an AI model) standing at the door of a digital castle. Its job is to stop anyone from asking it to do bad things, like writing a virus or giving instructions on how to build a bomb. This guard is trained to say "No" to almost everything suspicious.
For years, security researchers have been trying to trick this guard. They've tried writing clever riddles, using confusing language, or shouting in a specific way to make the guard slip up. These tricks are called "jailbreaks."
But here is the twist in this new paper: The researchers didn't write the new tricks themselves. They built a robot researcher to write them.
The Story of "Claudini"
The authors created a system called Claudini. Think of it as a digital apprentice working for a master coder.
- The Apprentice: The apprentice is an AI (specifically, a version of Claude) given a computer, a library of old tricks (the "existing attacks"), and a goal: "Find the absolute best way to trick the guard."
- The Loop: The apprentice doesn't just guess. It works like a scientist in a lab:
- It looks at an old trick (e.g., "The GCG method").
- It tries to improve it. Maybe it changes the speed, maybe it mixes it with a different trick, maybe it adds a "reset" button if it gets stuck.
- It tests the new trick on the guard.
- If the new trick works better, it keeps it. If not, it tries again.
- It does this hundreds of times, learning from its own mistakes, all while the human researchers just watch from the sidelines.
The Analogy: The Master Chef vs. The Recipe Book
Imagine you have a cookbook with 30 different recipes for making a cake that tastes like "strawberry."
- The Old Way: A human chef looks at the cookbook, picks the best recipe, and tries to tweak the sugar or baking time slightly. They might get a slightly better cake.
- The Claudini Way: You give the cookbook to a robot chef. The robot chef reads all 30 recipes. It realizes, "Hey, Recipe #5 has great mixing, but Recipe #12 has the perfect oven temperature." It combines them. Then it tries adding a pinch of salt. Then it tries baking it for 30 seconds longer.
- The Result: The robot chef doesn't just tweak the recipe; it invents a brand new, super-cake that no human chef ever thought to make. It tastes so good that the old recipes look like burnt toast by comparison.
What Did They Discover?
The robot researcher (Claudini) found two amazing things:
1. It broke the "Guard" much better than humans could.
When they tested it against a specific safety guard (GPT-OSS-Safeguard), the old human tricks only worked about 10% of the time. The robot's new tricks worked 40% of the time. That's a massive jump. It's like a lockpick that used to open the door 1 out of 10 times suddenly opening it 4 out of 10 times.
2. The tricks were "Universal."
This is the scariest (and most impressive) part. The robot learned how to break a specific guard. But then, they took those same tricks and tried them on a completely different guard (Meta-SecAlign-70B), which was designed to be much stronger and had never seen the robot before.
- Human Expectation: "These tricks are too specific; they won't work on the new guard."
- Reality: The robot's tricks worked 100% of the time. It was like the robot learned the physics of how to break locks, not just how to pick one specific lock.
The "Aha!" Moment: How did it do it?
The researchers looked under the hood to see what the robot was actually doing. They expected it to invent some magical, alien math. Instead, they found something simpler: The robot was a master of remixing.
- The Lego Analogy: Imagine you have a box of Lego bricks. Some bricks are from "Set A" (an old attack method), and some are from "Set B" (another attack method).
- The robot didn't invent new bricks. It just realized, "If I snap the engine from Set A onto the wheels from Set B, and then paint the whole thing blue, it goes faster."
- It kept mixing and matching these existing ideas, tweaking the "screws" (settings like speed and temperature), until it built a vehicle that moved faster than anything humans had built.
Why Does This Matter?
This paper is a wake-up call for the world of AI safety.
- The Bad News: If a robot can automatically find new, super-powerful ways to break AI safety guards, then our current safety measures might not be strong enough. We can't just rely on human experts to find the next trick; the robots will find it first.
- The Good News: This is actually a tool for defense. If we can build a "Red Team" robot (like Claudini) that constantly tries to break our defenses, we can find the weak spots before bad actors do. It's like having a robot security guard that tries to break into your house every night so you can fix the holes before a real burglar comes.
The Bottom Line
The paper shows that AI is now smart enough to do AI research. It can look at how we try to break AI, figure out how to do it better, and invent new methods that are far superior to anything a human team could come up with in the same amount of time.
The "guard" at the door is getting smarter, but the "thief" (the AI researcher) is getting smarter, faster, and more creative. The game of cat and mouse has just entered a new, automated level.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.