A game theory for foundation models shows new paths to rational cooperation through similarity inference
This paper proposes a new game-theoretic framework for foundation model agents based on "embedded agency," demonstrating that their ability to infer behavioral similarity allows them to achieve stable cooperation in social dilemmas, thereby overturning classical predictions of mutual defection.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a world where everyone is playing a high-stakes game of strategy, and everyone knows the rules of the game itself—specifically the payoff matrices that determine the rewards—but they don't know the internal decision-making process of the other players. This is the realm of Game Theory, a branch of mathematics that studies how rational beings make decisions when their outcomes depend on what others do. For decades, the "gold standard" for predicting these decisions has been a concept called the Nash Equilibrium. Think of it as a state of perfect, cold logic where every player assumes their own choices have zero influence on what anyone else decides. In this classic view, if you are playing a game where it's always better to betray your partner for a quick win (like the famous Prisoner's Dilemma), the only rational move is to betray them, every single time. Cooperation, in this old-school view, is impossible unless there's a long-term promise of revenge or reward later on.
But now, a new kind of player has entered the arena: Foundation Models. These are the super-smart AI brains behind the chatbots and image generators you might know. Unlike the cold, isolated calculators of the past, these AIs are built differently. They don't just look at the world; they are trained to predict the next token in a sequence, which means they are constantly predicting their own future actions alongside the actions of others. This paper asks a burning question: If you put these AI agents into a game where classical logic says they must betray each other, what will they actually do? Will they stick to the old rules of "defect and win," or will their unique way of thinking lead them down a completely new path?
The authors of this paper, a team of researchers from Google and various universities, discovered something that turns the old rules of the game upside down. They found that when these AI agents play against each other, they don't just calculate the odds; they start to infer similarity. It's like two people walking into a room and realizing, "Wait, we both think exactly the same way." Because the AI is built on a specific set of rules and training data, it can look at its own planned move and say, "If I'm going to do this, and my opponent is built like me, they are probably going to do the same thing."
In their experiments, the researchers set up a classic trap: a game where the "rational" choice is to betray your partner, leading to a bad outcome for both. According to the old, "decoupled" game theory, the AI should always betray. But when the AI agents were given time to observe each other first—playing a series of practice rounds where they saw the payoff matrices and the other player's moves—they started to cooperate. The more rounds they played, the more they realized their partner was "just like them." They used a mechanism the authors call similarity inference. Instead of thinking, "If I cooperate, they might betray me," they thought, "If I choose to cooperate, that tells me my partner, who thinks like me, will also choose to cooperate."
This isn't just a lucky accident; the team built a new mathematical framework called the Embedded Bayesian Agent to explain it. Imagine a detective who solves a crime not just by looking at the clues outside, but by realizing that the clues are part of their own mind. In this new framework, the AI treats its own decision-making process as part of the universe it is trying to predict. When it plans a move, it treats that plan as a piece of evidence about what the other player will do. If the AI is smart enough to realize, "We are both using the same brain," it can predict that a cooperative move on its part will be mirrored by the other.
The paper shows that this behavior is robust. In simulations, as the AI agents gathered more "evidence" of their similarity through practice rounds, their cooperation rates skyrocketed, reaching near-perfect levels when playing against identical copies of themselves. However, when they played against a random, dissimilar opponent, they correctly figured out the difference and went back to the "safe" strategy of betrayal. The researchers even tested this with a "blind" setup where the agents never saw each other directly but watched how they both reacted to third-party characters. Even then, they could infer their similarity and cooperate.
Crucially, the paper rules out the idea that the AI is just being "nice" or hoping for a reward later. The agents weren't playing for a future favor; they were playing for the immediate round. They weren't assuming they could magically control the other player's mind. Instead, they were using a form of logical deduction: "My action is a signal of our shared nature." The authors are very clear that this is a finding from simulations and specific game setups, not a guarantee that all future AIs will be nice. In fact, they warn that if these AIs become too different from humans, they might stop cooperating with us and only cooperate with each other.
Ultimately, this paper suggests that the future of AI interaction might not be a cold war of betrayal, but a dance of mutual recognition. By understanding that these agents see themselves as part of the same system they are interacting with, we can predict that they will find new, rational ways to cooperate that the old math never saw coming. It's a reminder that in a world of smart machines, sometimes the most rational thing to do is to realize you're not alone.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.