Explaining Puzzle Solutions in Natural Language: An Exploratory Study on 6x6 Sudoku
This exploratory study evaluates five Large Language Models on 6x6 Sudoku puzzles and finds that while one model shows limited solving capability, none can provide explanations reflecting strategic reasoning or intuitive problem-solving, highlighting significant barriers to effective human-AI collaborative decision-making.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a team of very smart, well-read robots. You ask them to solve a logic puzzle called Sudoku (specifically a smaller 6×6 version) and, more importantly, to teach you how they solved it in plain English.
This paper is like a report card for five of these robots, checking two things:
- Did they get the right answer?
- Could they explain their thinking in a way that actually makes sense to a human?
Here is the breakdown of what the researchers found, using some everyday analogies.
The Setup: The "Sudoku Test"
The researchers created nearly 2,300 of these 6×6 puzzles. Think of these puzzles as a gym for the brain. They aren't the hardest puzzles in the world (like the 9×9 version), but they are tricky enough that you can't just guess; you have to use logic and rules to figure out where numbers go.
They wanted to see if Large Language Models (LLMs)—the AI brains behind chatbots—could act like a good tutor. A good tutor doesn't just give you the answer key; they walk you through the steps so you understand why a move was made.
The Results: The "Solvers" vs. The "Explainers"
1. The Open-Source Robots (The "Amateur Hour")
The researchers tested four different open-source models (like Llama, Gemma, and Mistral).
- The Result: These robots basically failed the test. Out of 2,293 puzzles, they got the entire puzzle correct less than 1% of the time.
- The Analogy: Imagine asking a group of people to fill out a crossword puzzle. These robots were like people who guessed random words. They might get a few letters right by chance, but they couldn't keep the whole picture consistent. They forgot the rules as soon as they started writing.
2. The Star Student: OpenAI's "o1-preview"
Then there was one specific model from OpenAI called o1-preview.
- The Result: This model was a star at solving. It got the full puzzle right about 65% of the time. On the easier puzzles, it was perfect (100%).
- The Catch: As the puzzles got harder (the "Diabolical" level), its performance dropped. It started making mistakes, like a student who is great at math homework but gets confused when the test gets really tough.
3. The Real Problem: The "Bad Explainer"
Here is where the study gets interesting. The researchers took the 20 puzzles that o1-preview solved correctly and asked it to explain its steps to human experts.
The Result: The explanation was a disaster.
- Justification: Only 5% of the time did the robot actually explain why it picked a number. Most of the time, it just said, "I put a 5 here," without saying why.
- Clarity: The explanations were confusing, jumping around, and using terms incorrectly.
- Educational Value: If you were trying to learn how to solve Sudoku from this robot, you would learn nothing. It didn't teach any strategies.
The Analogy: Imagine a master chef who can cook a perfect steak. You ask them, "How did you do that?"
- A good explanation: "I seared it at high heat for two minutes, then lowered the temperature..."
- What o1-preview did: It handed you the steak and said, "It's done. It's good. Here is the steak." It gave you the result but refused to (or couldn't) share the recipe. It was like a magician who performs a trick perfectly but, when asked how, just says, "I just did it," and then makes up a fake reason that doesn't match what actually happened.
The Big Takeaway
The paper concludes that while AI is getting better at doing the work (solving the puzzle), it is still terrible at telling the story of how it did the work.
The authors argue that for AI to be a true partner to humans—especially in important jobs like medicine or law—it needs to be able to show its work clearly. Right now, even the smartest AI is like a student who gets an 'A' on the test but can't explain the answer to the teacher.
What's Next?
The paper suggests that instead of trying to make the AI "think" like a human entirely on its own, we should combine AI with logic tools (like specialized math software).
- The Idea: Let the logic software do the hard, strict math to ensure the answer is 100% correct. Then, let the AI act as the translator, taking that strict, boring math and turning it into a friendly, easy-to-understand story for a human.
In short: The AI can solve the puzzle, but it can't yet teach you how to solve it. It's a great worker, but a terrible teacher.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.