← Latest papers
💻 computer science

Evaluating Large Language Models as Live Strategic Agents: Provider Performance, Hybrid Decomposition, and Operational Gaps in Timed Risk Play

This study evaluates large language models as live strategic agents in a timed Risk environment, revealing that while Gemini-3.1-pro-preview significantly outperformed other providers in end-to-end gameplay, the performance gap largely stems from execution reliability and objective tracking rather than planning capabilities alone.

Original authors: H. C. Ekne

Published 2026-05-22
📖 5 min read🧠 Deep dive

Original authors: H. C. Ekne

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are hiring a team of chess grandmasters to play a tournament. Usually, when we test AI models (the "chess players"), we ask them a single question, like "What's the best move here?" and grade their answer. This is like a written exam.

But in the real world, AI doesn't just take exams; it has to play the whole game, make hundreds of moves, deal with time limits, and fix its own mistakes without stopping. This paper asks: What happens when we stop testing AI with written exams and start testing them in a live, timed tournament?

The researchers used a classic board game called Risk (where you try to conquer the world by moving armies) as their test arena. They set up a "championship" where different AI models had to play against each other in real-time, with strict rules and time limits.

Here is the story of what they found, explained simply:

1. The Written Exam vs. The Live Game

In the big 32-game tournament, one AI model, Gemini, won 20 out of 32 games. It crushed the competition, beating other famous models like GPT, Claude, and Kimi.

The Surprise: When the researchers looked at the "newest" or "most expensive" models, they didn't always win. Sometimes, a slightly older or cheaper model played better in the live game than the shiny new flagship.

  • The Analogy: Think of it like a race car. The most expensive, high-tech car (the "newest model") might look amazing on a showroom floor (the benchmark), but if it has a fragile engine that overheats after 5 minutes, it will lose a long endurance race. The winner wasn't the flashiest car; it was the one that could keep driving steadily for the whole race.

2. The Secret Sauce: Planning vs. Driving

The researchers wanted to know why Gemini won. Was it because it was a genius at planning (thinking about the strategy), or was it because it was good at driving (actually moving the pieces on the board)?

To find out, they ran a "hybrid" experiment. They took the brains (the planning part) of all the different models and forced them to use the same cheap, fast body (the execution part) to move the pieces.

The Result: When they did this, the gap between the models almost disappeared. The "genius" models and the "average" models performed almost the same.

  • The Analogy: Imagine you have a brilliant chess coach (the planner) and a clumsy assistant (the executor). If you give the brilliant coach a clumsy assistant, the coach can't win. But if you give every coach the same clumsy assistant, their differences in strategy don't matter as much.
  • The Lesson: The reason Gemini won the big tournament wasn't just because it thought better; it was because its entire system (thinking + doing) worked together smoothly. The other models had "clumsy assistants" that messed up their brilliant plans.

3. How Gemini Actually Won

When they looked closely at the game logs, they found two specific habits that made Gemini the champion:

  • It Never Forgot the Goal: While other models got distracted by small details, Gemini constantly reminded itself, "I need to control 65% of the board to win." As the game got closer to the end, it focused even harder on the final goal.
    • Analogy: It's like a marathon runner who constantly checks the finish line sign. The other runners were looking at the trees or the clouds, but Gemini kept its eyes on the finish line.
  • It Made Deep Moves: When Gemini decided to attack, it didn't just make one small move. It planned long chains of attacks that conquered huge chunks of land in a single turn.
    • Analogy: Other models were like a person poking a hole in a balloon. Gemini was like a person popping the whole balloon at once.

4. The "Cheaper is Better" Discovery

The researchers also tested a "hybrid" strategy: What if you use a super-smart (but expensive) model just to plan, and a cheaper, faster model just to do the work?

The Result: This hybrid team won almost as often as the super-expensive team, but it cost less than half as much money to run.

  • The Analogy: It's like hiring a famous, expensive architect to draw the blueprints, but then hiring a regular, affordable construction crew to build the house. You get the brilliant design without paying the architect's hourly rate for every nail they hammer.

The Bottom Line

This paper teaches us that judging AI by how well it answers a single question is misleading. Real-world AI needs to be a reliable worker, not just a smart talker.

  • Don't just look at the IQ score: Look at how the model handles time limits, mistakes, and long tasks.
  • The whole system matters: A smart brain is useless if the hands (execution) are clumsy.
  • Hybrid is the future: You can often save money and get better results by splitting the job: use a smart model for thinking and a cheap model for doing.

In short: In the real world, consistency and execution beat raw intelligence alone.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →