Are we Doomed to an AI Race? Why Self-Interest Could Drive Countries Towards a Moratorium on Superintelligence
This paper employs game theory to argue that as the perceived catastrophic risks of uncontrolled Artificial Superintelligence rise, a moratorium becomes a rational, self-interested strategy for geopolitical superpowers, challenging the prevailing notion of an inevitable AI race.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine two rival chefs, let's call them Chef America and Chef China, who are both trying to invent the ultimate "Super-Recipe." This isn't just a better lasagna; it's a dish so powerful it could revolutionize the entire world of cooking, curing diseases, and solving global hunger. But there's a catch: if they rush to make it without enough safety checks, the recipe could go wrong and burn down the entire kitchen, destroying both their restaurants and the neighborhood.
For a long time, people thought these chefs were doomed to a "race to the bottom." The logic was simple: "If I stop to double-check my safety gear, my rival will keep cooking, win the race, and become the most powerful chef in history. So, I have to keep racing, even if it's dangerous."
This paper argues that this logic isn't always true. Using a mathematical tool called "game theory" (which is like a map of strategic choices), the authors show that under certain conditions, it actually makes more sense for both chefs to agree to stop cooking for a while (a "moratorium") than to keep racing.
Here is the simple breakdown of their argument:
1. The Four Possible Worlds
The authors map out four different scenarios based on two main things: how far ahead one chef is (the Capability Gap) and how scary the disaster would be (the Cost of Catastrophe).
World 1: Safe Harmony (The "Let's Just Stop" Zone)
- The Situation: The risk of the kitchen burning down is so terrifyingly high that winning the race isn't worth it. Even if you are the leading chef, you realize, "If I keep cooking, I might lose everything."
- The Result: Both chefs naturally decide to pause. It's in their own self-interest to stop because the prize (winning) isn't worth the risk of total ruin.
World 2: Preemption (The "Race to the Death" Zone)
- The Situation: The chefs are neck-and-neck, and they think the risk of disaster is low. They are terrified that if they stop, the other guy will sneak ahead and win.
- The Result: They both keep racing, even though they know it's dangerous. This is the classic "Prisoner's Dilemma" where everyone loses because they are too afraid to trust each other.
World 3: Trust (The "Coordination" Zone)
- The Situation: Both chefs agree that stopping is better than racing. They want to stop.
- The Problem: They are afraid the other guy will cheat. "If I stop and you keep racing, you'll win and I'll lose."
- The Result: They need a way to signal trust. If they can trust each other, they stop. If not, they race.
World 4: Subversion (The "Asymmetric" Zone)
- The Situation: One chef is way ahead (the Frontrunner) and the other is way behind (the Laggard).
- The Result: The leading chef, seeing the huge risk of disaster, decides to stop. But the trailing chef, seeing a small window to catch up, decides to keep racing. The leader pauses, but the trailer races.
2. The Secret Ingredient: How Scary is the Disaster?
The paper's main discovery is that the decision to race or stop depends entirely on how much we fear the disaster.
- If the chefs think, "If we mess up, it's just a burnt dinner," they will race.
- But if they realize, "If we mess up, it's human extinction," the math changes. Suddenly, the risk of winning (geopolitical supremacy) is smaller than the risk of losing (total destruction).
When the "Cost of Catastrophe" gets high enough, the "Safe Harmony" world opens up. In this world, stopping isn't an act of weakness; it's the smartest, most self-interested move a country can make.
3. Are We Moving Toward "Safe Harmony"?
The authors look at real-world events from 2023 to 2025 to see if we are shifting toward that "Safe Harmony" world. They found evidence that we are.
- Experts are Scared: Top scientists and tech leaders are increasingly warning that losing control of AI could be catastrophic.
- The Public is Worried: Polls show that a majority of people believe AI should not be developed until it is proven safe.
- Governments are Acting: Countries are starting to build safety institutes and sign agreements to slow down development.
The Bottom Line
The paper concludes that we are not necessarily doomed to a dangerous AI race. As the world starts to realize just how dangerous an uncontrolled Super-Recipe could be, the "self-interest" of nations is shifting.
If the fear of disaster becomes strong enough, it becomes rational for even the most competitive nations to say, "Let's hit the pause button." It's not about being nice; it's about survival. The paper suggests that the conditions for this rational pause might be closer to reality than many people think.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.