← Latest papers
📈 economics

Falling Behind Drives Unsafe Development in an Idealised AI Race Experiment

This paper combines a behavioral experiment and an evolutionary model to demonstrate that unsafe AI development is driven less by individual risk preferences and more by competitive dynamics—specifically the fear of falling behind and reactive responses to opponents—suggesting that policy should prioritize reducing competitive pressure and fostering cooperation over focusing solely on individual risk management.

Original authors: Elias Fernández Domingos, The Anh Han

Published 2026-07-29
📖 4 min read☕ Coffee break read

Original authors: Elias Fernández Domingos, The Anh Han

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a high-stakes game of "Red Light, Green Light," but instead of a doll, you're racing against a rival to build the most powerful robot in the world. In the real world, this is what happens when companies or countries compete to create Artificial Intelligence (AI). They are in a frantic race to be the first to finish, because being first often means winning the biggest prizes, like market dominance or national security. But here's the tricky part: taking shortcuts to go faster can be dangerous. If you rush, you might skip safety checks, which could lead to a catastrophic accident later. This creates a tense tug-of-war: do you play it safe and risk losing the race, or do you speed up and risk blowing everything up? Scientists call this the "AI Race," and they've been wondering: do people choose to take these dangerous shortcuts because they are naturally reckless, or because the pressure of the race itself forces their hand?

To find out, researchers Elias Fernández Domingos and The Anh Han set up a digital playground. They paired up strangers and put them in a simulated AI race. In this game, players could choose between "Safe" development (which moves them forward slowly but steadily) or "Unsafe" development (which zips them ahead quickly but adds a hidden "risk meter" to their score). If a player won the race while their risk meter was too high, there was a chance they would lose their prize. The researchers changed how high the risk meter could go in different groups of players—some had a low cap (10% chance of disaster), some medium (60%), and some high (90%). They wanted to see if a higher risk cap would make people play safer, or if the fear of losing to the other player would make them play unsafe anyway.

The results were a surprise. The researchers had guessed that people with a higher risk cap would be more careful, and that people who naturally liked taking risks would choose the unsafe path more often. But the data said "nope" to both of those ideas. It didn't matter if the risk cap was low or high; it didn't matter if a player was naturally brave or cautious. Instead, the players' choices were driven by the story of the race itself.

Think of it like a video game where you are chasing a rival. If you are ahead, you feel safe and start playing carefully, like a champion protecting their lead. But if you fall behind, panic sets in. The players who were losing were much more likely to choose the "Unsafe" option, hoping the speed boost would help them catch up. It was a "fear of falling behind" mechanism. Even more interesting, players watched their opponents like hawks. If the other person chose the risky shortcut, the player was significantly more likely to copy them in the next round. It wasn't about being reckless; it was about reacting. If your rival is speeding, staying slow feels like a losing strategy, so you speed up too, even if it's dangerous.

The researchers also noticed that the very first move a player made set the tone for the whole game. If someone started with a risky move, they tended to keep taking risks later on. To understand why this happens, the team built a computer model with four types of "AI personalities": one that is always safe, one that is always unsafe, one that starts safe and copies the opponent, and one that starts unsafe and copies the opponent. This model showed that in a competitive race, the "copycat" strategies often win out. The model suggests that unsafe behavior isn't just about individual bad apples; it's a chain reaction. One person takes a risk, the other feels forced to follow, and suddenly everyone is speeding toward a cliff, not because they love danger, but because the race structure makes safety feel like a losing move.

So, what does this mean for the real world? It suggests that we can't just rely on hoping AI developers are naturally careful. The pressure of the race itself is the real driver. If the competition makes it feel like "safety means losing," people will choose speed. To fix this, the paper suggests we need to change the rules of the race itself—perhaps by making it easier to cooperate, slowing down the competition, or ensuring that no one feels forced to take dangerous shortcuts just to catch up. The danger isn't just in the people; it's in the game they are forced to play.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →