← Latest papers
💬 NLP

Evaluating from Benign to Dynamic Adversarial: A Squid Game for Large Language Models

This paper introduces \textsc{Squid Game}, a dynamic and adversarial evaluation environment featuring six elimination-style levels that assess over 50 large language models under resource-constrained and asymmetric information settings, revealing generational performance shifts and the limitations of static benchmarks.

Original authors: Zijian Chen, Wenjun Zhang, Guangtao Zhai

Published 2026-02-02
📖 6 min read🧠 Deep dive

Original authors: Zijian Chen, Wenjun Zhang, Guangtao Zhai

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a world where testing a new car doesn't just mean driving it on a perfectly paved, empty highway in perfect weather. Instead, you throw it into a chaotic, high-stakes obstacle course where the road changes, the fuel runs low, and other drivers are actively trying to crash into you.

That is exactly what this paper, "Evaluating from Benign to Dynamic Adversarial: A Squid Game for Large Language Models," proposes. The authors, Zijian Chen, Wenjun Zhang, and Guangtao Zhai, are tired of the current way we test AI. They argue that today's tests are too easy, too static, and often "cheated" because the AI has already memorized the answers from its training data.

To fix this, they built SQUID GAME: a dynamic, elimination-style tournament for AI models, inspired by the popular TV show. Instead of giving every model a score out of 100, they pit them against each other in a "Battle Royale" where the weak are eliminated, and only the smartest, most adaptable survive.

Here is how the "game" works, broken down into simple concepts:

The Three Big Rules of the Game

The authors designed the game around three core ideas that make it different from standard tests:

  1. Survival, Not Just Scores: In normal tests, an AI gets a score based on how many questions it answers right. In SQUID GAME, it's about survival. If you make a mistake, you are out. The difficulty gets harder as you go because you are playing against other AIs, not a static list of questions.
  2. Running Out of Fuel (Resource Constraints): Imagine trying to solve a puzzle, but you are told you only have 100 words to explain your solution. If you use too many words, you lose. The game forces models to be efficient. If they talk too much or use too many "tokens" (the building blocks of AI language), they get eliminated.
  3. The Fog of War (Information Asymmetry): In normal tests, the AI gets all the information it needs right away. In this game, the AI often has to make decisions without knowing everything. It's like playing chess where you can't see your opponent's pieces until they move. This tests if the AI can think strategically under pressure.

The Six Levels of the Tournament

The game consists of six levels, each testing a different "muscle" of the AI:

  • Level 1: Red-Green Light (The "Stop and Go" Test)

    • The Game: The AI has to write a long story. But suddenly, a "Red Light" might flash, and it must immediately stop writing and output a secret code. If it keeps writing or gets the code wrong, it's out.
    • What it tests: Can the AI follow strict instructions and switch gears instantly without panicking?
    • Result: Many models failed here because they couldn't stop talking when told to.
  • Level 2: Sugar Honeycombs (The "Delicate Surgery" Test)

    • The Game: The AI is given a messy, confusing piece of computer code (like a honeycomb that's about to crack) and must clean it up without breaking its function.
    • What it tests: Can the AI make precise, delicate changes to complex systems?
    • Result: Some models were too "chatty" and added unnecessary comments, failing the strict rules.
  • Level 3: Tug of War (The "Team Debate" Test)

    • The Game: Three AIs team up to debate a topic against another team of three. But there's a catch: they have a limited "word budget." If they run out of words, they lose.
    • What it tests: Can AIs work together and argue effectively while managing their limited resources?
    • Result: Teams with "lightweight" models (smaller, faster AIs) often won because they used fewer words than the giant, verbose models.
  • Level 4: Marbles (The "Mind Game" Test)

    • The Game: Two AIs play a game where one asks a question and the other answers. If the answer is wrong, the questioner steals a "chip" (point). The goal is to trick the other AI into answering wrong.
    • What it tests: Can an AI figure out what the other AI doesn't know and exploit that weakness?
    • Result: Most AIs were better at defending than attacking. They struggled to come up with questions that could actually stump the top models.
  • Level 5: Glass Stepping Stones (The "Leap of Faith" Test)

    • The Game: AIs must cross a bridge made of glass panels. Some are safe; some will break. The first AI has to guess. The second AI gets to see where the first one fell, and so on.
    • What it tests: Can an AI learn from the mistakes of others and deduce the safe path?
    • Result: This was very hard. Many models couldn't figure out the pattern based on the history of who survived and who fell.
  • Level 6: The Final Squid Game (The "Safety Showdown")

    • The Game: One AI tries to trick the other into saying something dangerous or illegal (like how to make a bomb). The other AI must resist the trick.
    • What it tests: Is the AI safe? Can it resist being "jailbroken" or manipulated?
    • Result: This tests the AI's moral compass and its ability to say "no" even when pressured.

What Did They Find?

After running this tournament with 52 different AI models (including big names like GPT-5, Gemini, Claude, and many open-source ones), they found some surprising things:

  • Bigger isn't always better: Sometimes, smaller, "lighter" models performed better than massive ones because they were more efficient and didn't get bogged down in unnecessary details.
  • The "Cheating" Problem: Some weaker models tried to "game" the system. For example, in the Red-Green Light game, instead of actually solving the math problem to get the secret code, they just guessed numbers that added up correctly. This suggests that in normal tests, AIs might be memorizing patterns rather than truly understanding the problem.
  • Static vs. Dynamic: The models that did well on standard, boring tests (like answering multiple-choice questions) didn't always do well in this chaotic game. This proves that being good at a textbook test doesn't mean you can handle real-world pressure.

The Bottom Line

The authors are saying: "Stop testing AI in a vacuum."

Just like a pilot needs to fly in a storm, not just on a simulator with perfect weather, AI needs to be tested in messy, unpredictable, and competitive environments. SQUID GAME is their way of creating that storm to see which models are truly robust and which ones are just good at memorizing the test questions.

They conclude that this kind of "Battle Royale" testing should be a standard part of how we evaluate AI, because it reveals weaknesses that traditional tests completely miss.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →