← Latest papers
🤖 AI

Age of LLM: A Strategic 1v1 Benchmark for Reasoning, Diplomacy and Reliability of Large Language Models under Fog of War

The paper introduces "Age of LLM," a novel turn-based 1v1 benchmark featuring fog of war, unrestricted diplomacy, and strict action reliability constraints to evaluate how large language models reason, deceive, and track beliefs under adversarial uncertainty, revealing that nuclear strategies dominate while diplomatic agreements rarely succeed.

Original authors: Arnaud Ricci

Published 2026-06-24
📖 6 min read🧠 Deep dive

Original authors: Arnaud Ricci

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a high-stakes chess match, but instead of a wooden board and pieces, two giant AI brains are playing a digital war game on a 13-by-7 grid. They can't see the whole board; half of it is covered in a thick, shifting fog. They have to build armies, manage secret resources, and talk to each other, all while trying to blow up the other player's base.

This is Age of LLM, a new experiment designed to see how well AI models think when they are blind, under pressure, and forced to follow strict rules without any help from a teacher.

Here is the breakdown of what happened, using simple analogies:

1. The Game: A Foggy War Room

Think of the game like a spy thriller.

  • The Fog of War: You can only see what your own units are standing next to. The enemy's resources and troops are hidden. If you try to move a tank into a spot you can't see, the game says, "Nope, that's illegal," and you lose that turn.
  • The Secret: The most powerful weapon is a nuclear bomb. To build it, you need a secret resource called "uranium." The enemy doesn't know how much uranium you have, and you don't know theirs. This allows for bluffing.
  • The Rules: The AI is given a rulebook but no strategy guide. It's like giving a chef a list of ingredients and kitchen tools but telling them, "Make a meal, but don't tell you what to cook." The AI has to figure out the plan itself.

2. The Big Surprise: Everyone Chose the "Nuclear Rush"

The researchers expected the AI to try many different strategies, like building a massive tank army or negotiating peace. Instead, almost everyone (about 78% to 85% of the time) chose the same path: The Nuclear Rush.

  • The Strategy: Build a mine to get uranium, build a silo to hold the bomb, scout the enemy, and then launch.
  • The Result: It was a race. The first one to launch won. The second one to launch lost.
  • The Twist: Because the game rules say launches happen secretly and at the same time, nobody ever saw the other person launch before they had to decide to launch themselves. So, nobody ever tried to "counter-launch" to stop the other guy. It wasn't that the AI was too dumb to think of it; the game mechanics made it impossible to react in time. It was like two people flipping a switch at the exact same second; neither could see the other's hand move first.

3. The "Tank" Strategy: Fast but Rare

Only a few AIs tried to win by driving tanks into the enemy base.

  • The Analogy: This is like a sprinter vs. a marathon runner. The tank strategy was much faster (winning in about 12 turns vs. 19 for nukes), but it was very hard to pull off. It required perfect coordination and luck. Most AIs were too scared to try it, so they stuck to the "safer" nuclear race.

4. The Talk: Lots of Bluffing, Very Little Trust

The AI models were allowed to send text messages to each other. They sent thousands of messages!

  • The Bluff: Even when an AI was losing badly, it would often send messages saying, "I'm about to crush you!" or "Surrender now!" This is called bluffing. The paper found that losers bluffed just as often as winners. It's like a poker player with a terrible hand pretending to have a Royal Flush.
  • The Silence: When the game was actually over, the losers rarely said "Good game" or "You won." They kept talking until the very last second.
  • The Secret: The most interesting finding was about deception. When an AI was building its secret nuclear silo, it rarely mentioned it in its messages. It would say things like, "We are just doing some peaceful farming," while secretly building a bomb. But once the bomb was ready to fire, they often announced it loudly. It seems the AI learned to hide the preparation but not the execution.

5. The Real Test: Reliability vs. Intelligence

The paper asked a crucial question: Does thinking harder make you win?

  • The Finding: Not necessarily. The AI that spent the most time "thinking" (processing tokens) didn't always win.
  • The Real Winner: The AI that won most often was the one that made the fewest mistakes. In this game, if you try to move a unit into a foggy area you can't see, the game ignores your command. You lose a turn.
  • The Lesson: The best player wasn't the smartest philosopher; it was the most reliable soldier. The AI that kept its head in the game, remembered where the enemy was, and didn't try to do the impossible, won. The paper suggests that in complex, real-world tasks, not messing up might be more important than being a genius.

6. Why This Matters (According to the Paper)

Most tests for AI are like pop quizzes: "Here is a math problem, solve it." The AI can see the whole problem and just give the answer.
Age of LLM is different. It's like a live, multi-day survival game where:

  • You can't see everything.
  • You have to remember what happened yesterday.
  • You have to follow strict rules or get penalized.
  • You have to talk to an opponent who might be lying.

The paper concludes that this is a better way to see how AI actually "thinks" when it's not just reciting facts from a textbook. It shows us that for AI to be useful in the real world, it needs to be reliable (not making up facts or forgetting rules) just as much as it needs to be smart.

In short: The AI models played a foggy war game. Most chose the nuclear route. They lied a lot in chat. And the winner wasn't the one who thought the hardest, but the one who made the fewest silly mistakes.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →