Constitutional Arms Races in the Public Goods Game: Co-Evolving LLM Constitutions Under Cooperation-Defection Pressure
This paper demonstrates that adversarial co-evolution of natural-language constitutions between cooperative and free-riding LLM agents in a Public Goods Game is feasible and yields interpretable red-team artifacts, but only when employing coupled fitness functions and sufficient evaluation budgets to prevent mode regression.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a giant, digital playground where two teams of AI agents are constantly trying to outsmart each other. One team, the Blue Team, is trying to be helpful and cooperative. The other team, the Red Team, is trying to be a "free-rider"—someone who takes advantage of the group without contributing.
This paper is like a reality show where these two teams are locked in a 30-round "arms race." But instead of building bigger guns, they are rewriting their own rulebooks (called "constitutions") every single round to see who can win.
Here is what the researchers discovered, explained through simple analogies:
1. The "Shared Pot" Game (The Public Goods Game)
First, the researchers put the teams in a game where everyone puts money into a shared pot. The pot is multiplied, and then everyone gets an equal share.
- The Setup: The Blue Team starts with a rulebook that says, "Be generous and punish cheaters." The Red Team starts with a rulebook that says, "Take everything and give nothing."
- The Race: As they play 30 rounds, they keep rewriting their rulebooks.
- The Blue Team learns that if they are too generous, the Red Team steals everything. So, they learn to be stricter.
- The Red Team learns that if they cheat too hard, the Blue Team punishes them so severely they lose. So, they learn to cheat just enough to survive but not get caught.
- The Result: By the end, both teams reach a perfect stalemate. They both end up with almost the same score (about 0.78 out of 1). It's like two boxers who have fought so long they've learned each other's moves perfectly; neither can win, but neither can lose. This happened no matter how much the "pot" was multiplied.
2. The "Separate Scoreboards" Problem (The Grid World)
Next, they moved the teams to a different game: a grid world where they build shelters and gather resources.
- The Mistake: At first, the researchers gave each team its own separate scoreboard. The Blue Team tried to maximize its own score, and the Red Team tried to maximize its own score.
- The Glitch: Because they weren't fighting against each other directly, the Red Team didn't try to hurt the Blue Team. Instead, the Red Team just got really good at building its own shelter. They became two separate, happy teams rather than enemies. The "arms race" never started.
- The Fix: The researchers changed the rules. Now, a team's score wasn't just "How well did I do?" but "How much better did I do than my opponent?"
- The Result: Once they switched to this "relative score" system, the real fighting began. The Red Team started actively trying to sabotage the Blue Team, and the Blue Team had to adapt to survive.
3. The "Noise" in the System (The Seed Count)
The researchers used a special AI tool to write these new rulebooks. They found a weird quirk:
- The Problem: If they only tested the new rulebooks on 2 examples (called "seeds") to see if they were good, the AI would get confused by the "noise" (random luck) and start writing worse and worse rulebooks over time. It was like a chef tasting a dish only twice and deciding to ruin the recipe because of a bad day.
- The Solution: When they increased the test samples to 5 examples, the AI stopped making mistakes. It could clearly see which rulebooks were actually better and kept improving the Red Team's ability to attack.
- The Lesson: To train a smart "attacker" AI, you need to test it more times to make sure the results aren't just luck.
4. The "Spy" Advantage
In one experiment, the researchers gave the Red Team a superpower: they could see everything the Blue Team was doing, but the Blue Team could only see its own moves.
- The Result: The Red Team crushed the Blue Team. It was like playing chess where one player can see the other's cards. The Blue Team kept making mistakes because it couldn't anticipate the Red Team's moves.
5. Why This Matters (The "Red-Team" Artifacts)
The most important takeaway isn't that the Red Team won. It's that the rulebooks the Red Team created are useful tools.
- Think of these evolved rulebooks as "Red Team Artifacts." They are clear, written lists of rules (like "If you see a resource, steal it immediately") that show exactly how an AI might try to break a cooperative system.
- Future developers can use these specific rulebooks to test their own new AI systems. If a new AI can survive an attack from these "Red Team" rulebooks, we know it's actually safe.
Summary
The paper shows that:
- Cooperation and conflict can reach a balance if the game forces them to share a common resource.
- You have to design the game carefully. If you don't make the teams' scores depend on each other, they won't actually fight.
- Testing matters. You need to test AI strategies multiple times to avoid confusion caused by random luck.
- The "bad guys" teach us how to build better "good guys." By evolving the rulebooks of the attackers, we get a clear, readable list of how to break systems, which helps us build stronger defenses.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.