Hallucinations on the Board: Tool-Augmented Evaluation of LLM Chess Commentary
This paper introduces ACT-Eval, a tool-augmented evaluation framework that decomposes chess commentary into atomic claims to assess factual correctness and strategic coverage, revealing that while tool integration reduces hallucinations in large language models, significant gaps in expert-level strategic explanation persist across both proprietary and open-weight models.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a robot how to play a game of chess. You have two powerful tools at your disposal. First, you have a "Super-Engine," a machine that can calculate millions of moves per second and knows the mathematically perfect move for any situation. It's like a genius mathematician who can solve a puzzle instantly but speaks only in cold, hard numbers. Second, you have a "Storyteller," a large language model (LLM). This is a robot that is incredibly good at writing, speaking, and explaining things in a way humans understand. It can tell a thrilling story about a battle, but it sometimes makes things up because it doesn't actually "see" the chessboard; it just guesses what words usually go together.
The big question scientists are asking is: Can we combine these two? Can we let the Storyteller use the Super-Engine's brain to write a commentary that is both exciting and 100% true? The problem is that the Storyteller often "hallucinates"—it confidently describes pieces on squares where they don't exist or invents moves that are illegal. Until now, we've had a hard time catching these lies. If you ask a robot to grade another robot's story, the grader might just be fooled by the fancy words, missing the factual errors entirely. It's like asking a student who doesn't know the rules of chess to grade a story about a chess game; they might give an "A" for good writing even if the student described a knight flying like a bird.
This paper, titled "Hallucinations on the Board," introduces a new way to test these AI storytellers called ACT-Eval. The researchers built a system that acts like a strict referee with a checklist. Instead of just reading the story and giving a score, ACT-Eval breaks the commentary down into tiny, individual facts—like "The Queen is on square E4" or "This move threatens a checkmate." Then, it sends each tiny fact to the Super-Engine to verify if it's actually true on the board. If the story says a piece is attacking a square, the system checks the board to see if it really is. If the story misses a key strategic idea that a human expert would spot, the system flags that too.
The results are a bit of a wake-up call. The researchers found that even the smartest AI models, when left to their own devices without the Super-Engine's help, get the facts wrong surprisingly often. For example, one top-tier model made factual errors in about 22% of its claims, while smaller models got it wrong more than 40% of the time. They would confidently say a piece was in a spot where it wasn't, or describe a threat that didn't exist. However, when the researchers gave these models access to the Super-Engine tools to check their facts before writing, the error rate dropped significantly. The models became much better at stating the truth.
But there's a catch. While the tools helped the models stop lying about the facts, they didn't necessarily make the models better at understanding the story of the game. Even with the tools, the AI models often missed the deep, clever strategic ideas that a human grandmaster would notice. They could tell you the pieces were in the right place, but they struggled to explain why the move was brilliant in a way that felt complete. The paper suggests that while we can fix the AI's ability to check the facts, teaching it to truly understand the "why" behind a move is still a very hard challenge. The study concludes that we need these tool-assisted checkers to trust what the AI says, but we shouldn't expect them to replace human experts just yet.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.