RevengeBench: Reverse Engineering Code-Space Policies from Behavioral Experiments
This paper introduces RevengeBench, a benchmark that evaluates the ability of large language models to reverse-engineer executable code policies from behavioral traces in game environments, demonstrating that targeted experimental design significantly improves reconstruction accuracy and yields competitive advantages in downstream tournaments.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to figure out how a master chef cooks a secret dish. You can't see inside the kitchen, and the chef won't tell you the recipe. All you can do is watch the chef cook the dish over and over again while they face different ingredients and challenges.
RevengeBench is a new "detective training ground" for Artificial Intelligence (AI). It asks a simple but tricky question: If an AI only watches another AI play a game, can it figure out the exact code (the "recipe") the other AI is using to make its moves?
Here is how the paper breaks it down, using everyday analogies:
1. The Setup: The "Black Box" Chef
In the past, scientists studying behavior (like animal behavior) could only guess what was happening inside an animal's brain by watching what the animal did. They couldn't look inside.
- The Analogy: Imagine a chef who never speaks. You only see them chop vegetables, stir pots, and plate food. You have to guess the recipe just by watching.
- The Twist: In this new benchmark, the "chef" is an AI playing a video game (like a snake game, a poker game, or a robot battle). The "detective" is another AI trying to write the exact code that makes the chef play that way.
2. The Two Ways to Investigate
The paper tests two different ways the detective AI can learn:
- Passive Observation (Just Watching): The detective AI sits back and watches the chef play against random opponents. It tries to guess the recipe based on what it sees.
- The Analogy: You sit in the dining room and watch the chef cook 20 times. You try to write down the recipe based on that.
- Active Intervention (The "Probe"): The detective AI gets to design its own "opponent" to play against the chef. It creates a specific situation to trick the chef into revealing its secrets.
- The Analogy: You realize the chef always avoids spicy peppers. So, you send in a waiter who only brings spicy peppers. If the chef suddenly changes their cooking style to avoid the heat, you learn something new about their rules.
- The Paper's Finding: Being able to "poke" the chef with these custom opponents helps the detective AI learn better, but only if the detective is smart enough to design a good poke. If the detective is confused, poking doesn't help.
3. The Results: How Good Are the Detectives?
The researchers tested 12 different "super-smart" AI models (the detectives) on 75 different "chefs" (hidden game strategies).
- The Score: They measured how close the detective's guessed recipe was to the real one.
- The best detectives could figure out about 72% of the chef's secret moves.
- The weakest detectives only figured out about 34%.
- The "Aha!" Moment: Even when the detective didn't get the recipe 100% perfect, the "rough draft" recipe they wrote was still useful.
- The Analogy: If you guess the chef uses "a lot of salt" instead of "exactly 3 teaspoons," you might not get the dish perfect, but you can still cook a meal that beats the chef in a contest. The paper found that weaker AI models, which usually struggle to beat the chef, suddenly got much better at winning once they had this "rough draft" of the opponent's code.
4. Why This Matters (According to the Paper)
This isn't just about winning video games. The paper suggests this is a way to test how well AI can understand how other AIs think.
- It's a Reverse Engineering Test: It shows that AI can look at behavior and work backward to find the hidden rules (the code) causing that behavior.
- It's Not Magic: The paper admits that sometimes the AI guesses wrong, or gets stuck in a loop of bad guesses. Also, some games (like the robot battle game) were much harder to figure out than others (like the poker game).
- The Bottom Line: You don't need to know the exact secret code to beat an opponent. You just need a "good enough" guess of how they think.
Summary in One Sentence
RevengeBench is a test where AI detectives try to reverse-engineer the secret code of other game-playing AIs just by watching them play, proving that even a "rough guess" of an opponent's strategy can help you win the game.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.