← Latest papers
🤖 AI

Humans Are More Diverse: Frontier LLMs Show Extreme Policies in Idealised AI Development Races

This paper demonstrates that frontier LLMs exhibit highly diverse and often non-strategic behaviors in simulated AI development races, revealing that strong rule recall does not guarantee accurate state tracking or payoff calculation, thereby necessitating rigorous validity checks and trajectory-level analysis before interpreting their actions as human-like or safety-aware.

Original authors: Phu Hoa Pham, Duy Minh Dao Sy, Trung Kiet Huynh, Phu Quy Nguyen Lam, Chi Nguyen Tran, Minh Trung Le, Phong Hao Le, Dinh Nam Nguyen, Thien Ky Nguyen Dong, Elias Fernandez Domingos, Le Hong Trang, The A
Published 2026-08-04
📖 4 min read☕ Coffee break read

Original authors: Phu Hoa Pham, Duy Minh Dao Sy, Trung Kiet Huynh, Phu Quy Nguyen Lam, Chi Nguyen Tran, Minh Trung Le, Phong Hao Le, Dinh Nam Nguyen, Thien Ky Nguyen Dong, Elias Fernandez Domingos, Le Hong Trang, The Anh Han

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a high-stakes game of "chicken" played not with cars, but with the future of artificial intelligence. In this scenario, multiple companies are racing to build the smartest AI. They face a tricky choice every turn: play it safe and move slowly, or take a risky shortcut to move faster. If everyone plays safe, they all get a decent reward. If one takes a risk while others play safe, the risk-taker zooms ahead and wins big. But if everyone takes risks, the race becomes dangerous, and the winner might lose everything due to a catastrophic accident. This is the "AI safety dilemma," a concept researchers use to understand why companies might rush dangerous technology even when they know it's risky. To study this, scientists often use "Large Language Models" (LLMs)—the same kind of smart computer programs that can write stories or answer questions—to act as the players in these races. The big question is: Do these AI agents actually understand the game and make smart, strategic choices like humans do, or are they just guessing based on how the questions are worded?

This paper, titled "Humans Are More Diverse: Frontier LLMs Show Extreme Policies in Idealised AI Development Races," dives deep into that question. The researchers set up a digital race where AI agents compete against each other, sometimes in pairs and sometimes in groups of up to five. Before they could trust the AI's behavior, they built a strict "audit gate." Think of this like a driving test: before you can race on the track, you have to prove you know the rules, can read the speedometer, and can calculate your fuel. The study found that while the AI models were excellent at memorizing the rulebook (recalling rules), they were surprisingly bad at keeping track of the race as it happened (tracking the state). They often forgot who was winning or how much risk they had accumulated, even though they could recite the rules perfectly.

When the researchers looked at how the AI actually played, they discovered something fascinating and a bit worrying. Unlike humans, who show a wide variety of strategies—some cautious, some aggressive, some changing their minds based on what their opponent did—the AI models were surprisingly rigid. If you asked one specific AI model to play, it would stick to one narrow style of playing, like a robot that only knows how to drive in a straight line. For example, one model might almost always choose the safe option, while another would almost always choose the risky shortcut, regardless of the situation. In contrast, human players were a chaotic mix of all these styles. The study also found that the AI's behavior was incredibly sensitive to tiny changes in how the game was described. If the researchers swapped the words "Safe" and "Unsafe" for random letters like "P" and "Q," the AI's strategy would completely flip, even though the game mechanics stayed exactly the same. This suggests that the AI isn't truly "thinking" about the strategy; it's just reacting to the specific words on the screen.

The researchers also tested what happens when you tell the AI to act like a "risk-loving" or "risk-averse" character. In humans, your personality doesn't change just because someone tells you to act differently; your actual choices in a game are only weakly linked to your pre-game personality tests. But for the AI, giving it a "risk-loving" persona instruction made it immediately start taking huge risks, almost as if the instruction was a magic switch. This shows that the AI isn't displaying a deep, internal understanding of risk; it's just following the prompt like a very obedient, but easily confused, student.

Ultimately, the paper suggests that we can't just assume AI agents are making smart, human-like strategic decisions in these simulations. Just because an AI produces a sequence of moves that looks like a game plan doesn't mean it understands the game. The study warns that before we use these AI simulations to predict how real-world AI races will play out, we need to check if the AI can actually keep score and track the game state. Without these checks, we might be mistaking a confused robot following a script for a strategic genius, which could lead to very wrong conclusions about how to keep AI development safe. The findings are exploratory, meaning they are based on specific tests with specific models, but they highlight a crucial gap: AI can be great at reciting rules but terrible at applying them in a changing world.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →