Evaluating Language Models' Evaluations of Games
This paper introduces a formalism and dataset to evaluate AI systems' ability to assess board games, revealing that while reasoning models align better with human judgments than non-reasoning models, their fit to human data weakens as they approach game-theoretic optimality and they exhibit unpredictable resource usage.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are at a party where everyone is trying to solve a giant, complex puzzle. Usually, when we test Artificial Intelligence (AI), we ask the AI, "Can you solve this puzzle?" and we see how fast or accurately it can find the answer.
But this paper asks a different, more subtle question: "Can the AI look at the puzzle and tell us if it's even worth solving?"
The researchers wanted to see if AI models can act like a game critic or a referee, rather than just a player. They tested this by giving the AI 121 brand-new board games (games the AI had never seen before) and asking two specific questions about each one:
- The "Fairness" Question: If two people played this perfectly, who would win? Is the game balanced, or is one player doomed from the start?
- The "Fun" Question: Would people actually enjoy playing this game?
Here is what they found, broken down into simple concepts:
1. The "Smart" vs. The "Human"
The researchers compared three types of "thinkers":
- The Instant Guesser: Standard AI models that just spit out an answer immediately.
- The Deep Thinker: Newer AI models that take time to "think" (using a chain of thought) before answering.
- The Human: Real people who looked at the game rules and gave their gut feeling.
The Result: The "Instant Guessers" were terrible at this. They sounded confident but were often completely wrong, mostly because they were guessing based on patterns they saw in their training data rather than actually analyzing the new rules.
The "Deep Thinkers" were much better. They actually looked at the rules, simulated the game in their "mind," and gave answers that were closer to what humans thought. However, they weren't perfect.
2. The "Goldilocks" Problem (Too Smart is Bad)
Here is the most surprising finding. The researchers found a weird, non-straight line relationship between how "smart" the AI was and how much it agreed with humans.
- Too Dumb: The AI guesses randomly and disagrees with humans.
- Just Right: The AI thinks hard and agrees with humans.
- Too Perfect: As the AI gets extremely good at calculating the mathematically perfect outcome (like a super-computer chess engine), it starts to disagree with humans again.
The Analogy: Imagine a game of "Rock, Paper, Scissors."
- A human thinks, "I'll throw Rock because I feel lucky."
- A "Just Right" AI thinks, "Humans often pick Rock, so I should pick Paper."
- A "Too Perfect" AI thinks, "Statistically, the optimal move is Paper, so I will pick Paper."
The paper found that when the AI gets too obsessed with being mathematically perfect, it stops understanding what a human would actually do or feel. It becomes so rational that it loses its "human touch."
3. The "Fun" Factor is Hard to Measure
When the AI was asked about "Fairness" (math), the results were consistent. But when asked about "Fun," the results were all over the place.
The Analogy: Asking an AI to judge "Fairness" is like asking it to measure the length of a table with a ruler. There is a clear answer. Asking it to judge "Fun" is like asking it to measure the "happiness" of a song. It's subjective and messy.
The paper found that different AI models had very different ideas of what makes a game fun. Some thought a short game was fun; others thought a long, strategic game was fun. Because "fun" is hard to define, the AI models were inconsistent, and sometimes even the most advanced models couldn't agree with each other, let alone humans.
4. The "Energy Drink" Problem
Finally, the researchers looked at how much "brain power" (computing resources) the AI used to answer these questions.
The Analogy: Imagine you are asked to decide if a new recipe is good.
- Sometimes you just glance at the ingredients list (low effort).
- Sometimes you read the whole recipe, imagine the taste, and check the cooking time (high effort).
The paper found that the AI models were very unpredictable. Sometimes they used a tiny amount of effort to judge a complex game, and sometimes they used a massive amount of effort to judge a simple one. They didn't seem to have a good sense of "when to stop thinking." They didn't know how to be "resource-rational"—they didn't know when to save their energy.
Summary
This paper suggests that for AI to be a true partner to humans, it needs to do more than just solve problems. It needs to be able to evaluate problems.
- Standard AI is like a calculator: fast at math, but bad at understanding the "vibe."
- Reasoning AI is better: it can think through the rules and get closer to human opinions.
- The Catch: If the AI gets too focused on being mathematically perfect, it stops sounding like a human. And when it comes to subjective things like "fun," even the smartest AIs struggle to agree on what makes a game enjoyable.
The authors conclude that we need to teach AI not just how to solve, but how to decide what is worth solving, and how to use its brain power efficiently without overthinking or underthinking.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.