← Latest papers
🤖 machine learning

Position: Adversarial ML for LLMs Is Not Making Any Progress

This position paper argues that adversarial machine learning research has regressed in the era of large language models, as the field now tackles increasingly ill-defined, difficult, and hard-to-evaluate problems, raising concerns that another decade of work may yield no meaningful progress.

Original authors: Javier Rando, Jie Zhang, Nicholas Carlini, Florian Tramèr

Published 2026-06-03
📖 6 min read🧠 Deep dive

Original authors: Javier Rando, Jie Zhang, Nicholas Carlini, Florian Tramèr

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the field of Artificial Intelligence security as a game of "Locksmith vs. Burglar."

For the last decade, the locksmiths (security researchers) and burglars (adversaries) have been playing a very specific, well-defined game. They used to test simple locks on small, sturdy boxes. The rules were clear:

  • The Goal: The burglar tries to open the box without breaking the lock.
  • The Rules: The burglar can only wiggle the lock a tiny, specific amount (like shaking a key just a millimeter).
  • The Score: If the box opens, the burglar wins. If it stays shut, the locksmith wins.

This game was hard, but everyone knew exactly how to measure who was winning.

The Paper's Big Claim
Now, the game has changed. Instead of small boxes, we are trying to secure massive, open-ended "smart assistants" (Large Language Models or LLMs) that can talk about anything and do anything. The authors of this paper argue that the game has become broken. They claim that in this new era, the field is spinning its wheels and not actually making real progress because the rules of the game are now a mess.

Here is why, explained through simple analogies:

1. The Rules of the Game Are Gone (Defining the Problem)

In the old game, "winning" was easy to spot: Did the box open?
In the new game, the "box" is a chatbot. The goal isn't just to open it; it's to make the chatbot say something "bad."

  • The Analogy: Imagine asking a burglar to "steal something bad." What is "bad"? Is it stealing a candy bar? A car? A secret recipe?
  • The Problem: Because "bad" is subjective, researchers can't agree on what counts as a successful attack. One researcher might say, "He made the bot say a swear word, so he won!" Another might say, "But the swear word was harmless, so he lost." Without a clear scoreboard, you can't tell if the locksmiths are actually getting better at their job.

2. The Burglar Has Infinite Tools (Solving the Problem)

In the old game, the burglar was only allowed to wiggle the lock a tiny bit. This kept the game fair and measurable.
In the new game, the burglar can do anything.

  • The Analogy: The burglar can now wear a disguise, bribe the guard, pick the lock with a laser, or just talk the guard into opening the door.
  • The Problem: Because the burglar has infinite ways to attack, researchers can't write a computer program to find the "worst-case" attack. In fact, the paper notes that humans are currently better at breaking these models than computers are. The best "burglars" are now people using creativity and social engineering, not math. This makes it impossible to build a "perfect lock" because you can't test against every possible human trick.

3. The Scoreboard is Broken (Evaluating the Results)

In the old game, you could count exactly how many boxes were opened.
In the new game, we use other AI bots to judge if the chatbot said something "bad."

  • The Analogy: Imagine asking a robot to judge if a painting is "ugly." The robot might just say, "If it's not a picture of a cat, it's ugly." Or, the robot might be tricked by the burglar into saying, "No, that's actually beautiful."
  • The Problem: The tools we use to measure success are flawed. They are easily tricked, they don't understand human nuance, and they often give false alarms. Furthermore, the "locks" (the AI models) are often owned by big companies that update them secretly. If a company changes the lock overnight, researchers can't compare their results from last month to this month. It's like trying to compare two photos of a house, but the house keeps changing its color and shape between shots.

The Case Studies (The Evidence)

The paper looks at six specific areas where this "broken game" is happening:

  • Jailbreaks: Trying to trick a bot into saying "no." It's hard to define what "no" looks like when the bot is chatty.
  • Un-finetunable Models: Trying to make a bot so smart it can't be taught new, dangerous tricks. But if the burglar can retrain the bot, the game changes entirely.
  • Poisoning: Trying to sneak bad data into the bot's training. But the training data is so huge and messy (like a library with millions of books) that you can't tell which book caused the problem.
  • Prompt Injections: Tricking a bot into ignoring its rules. Since the bot can talk forever, the burglar can sneak in bad instructions over a long conversation, making it hard to spot.
  • Privacy: Trying to see if the bot "memorized" a specific person's data. But since the bot learned from the whole internet, it's impossible to know if it learned from that specific person or just from similar information elsewhere.
  • Unlearning: Trying to make the bot "forget" a specific topic (like how to build a bomb). But you can't just delete a fact from a brain; you might accidentally delete the bot's ability to do biology entirely.

The Conclusion

The authors aren't saying we should stop trying to secure AI. They are saying that we are trying to solve a puzzle with missing pieces and no picture on the box.

They argue that for the next decade, we might keep publishing papers that look like progress, but because the rules are fuzzy and the testing is unreliable, we aren't actually getting safer.

Their Advice:
Instead of trying to secure the entire, chaotic, open-ended AI, researchers should pick small, specific, well-defined "toy" problems (like the old "wiggle the lock a millimeter" game). If we can't even solve those small, clear problems with rigor, we have no hope of solving the massive, messy real-world problems.

In short: We are trying to build a fortress against an enemy that can change shape, attack from any angle, and judge its own victory, using a ruler that keeps shrinking. Until we fix the rules and the ruler, we aren't really making progress.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →