DungeonBench: A Benchmark for Rules-Rich Tactical Reasoning in Dungeons & Dragons Combat
The paper introduces DungeonBench, a comprehensive benchmark for evaluating tactical reasoning in Dungeons & Dragons combat that challenges frontier language models to navigate complex rules, geometry, and resource management across both single encounters and multi-encounter days, revealing significant gaps in their ability to balance immediate tactical advantages with long-term strategic sustainability.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a robot how to play a game. For a long time, scientists have been great at teaching robots to play games where the rules are simple and the goal is clear, like moving a piece across a grid or shooting a target. But there is a special kind of thinking that is much harder to teach: tactical reasoning. This isn't just about knowing the rules; it's about understanding how the rules bump into each other. It's the difference between knowing you can jump and knowing that if you jump now, you might land in a puddle that makes you slip, or that jumping later will let you grab a power-up before a timer runs out. This kind of thinking requires juggling geometry (where things are), timing (when things happen), and scarce resources (like limited energy or ammo). Scientists care about this because if we can teach computers to make these complex, "what-if" choices in a game, we get closer to building AI that can solve real-world problems where everything is connected and nothing is simple.
Enter DungeonBench, a new "test track" designed to see if artificial intelligence can actually master this kind of messy, rule-heavy thinking. The researchers built a digital simulator based on the famous tabletop game Dungeons & Dragons, but they stripped away the storytelling and the human referee. Instead, they created a strict, mathematical version of combat where every move, spell, and monster ability is a hard-coded rule. They didn't just ask the AI to win a single fight; they asked it to survive a whole "adventuring day" filled with multiple battles, where using too much magic in the first fight might leave the team helpless in the final one.
The paper tests some of the smartest AI models available today (like GPT-5.5 and Gemini 3.1 Pro) to see if they can act like a seasoned dungeon master. The setup is fair: the AI sees the entire battlefield, knows all the rules, and has a list of legal moves to choose from. The task is to pick the best move from hundreds of options, considering things like "If I cast this fireball here, will it hit the boss but also burn my friend?" or "Should I save my healing potion for now, or use it to finish this fight quickly?"
The results are a mix of impressive skill and funny failures. When the AI faced just one battle (the "Encounter" track), the top models were quite good, winning about 82% to 83% of the time. They were smart enough to position themselves well and use their spells effectively. However, when the researchers linked those battles together into a full day (the "Day" track), the AI's performance crashed. In this harder mode, the best models only cleared 40% of the days, and one model failed to clear a single day.
Why did they fail? The paper suggests that while these AIs are great at the "here and now," they struggle with resource budgeting. They tend to use their best spells and health potions to win the first easy fight, leaving them with empty hands and low health for the tough boss at the end of the day. They win the battle but lose the war. The researchers found that even though the AI could see the whole schedule of fights ahead of time, it couldn't figure out how to save its strength for later.
In short, DungeonBench shows that current AI is getting very good at following rules and making local decisions, but it still hasn't quite learned the art of long-term strategy. It's like a chess player who can win a single game by tricking the opponent, but who keeps running out of pieces before the tournament is over. The paper concludes that while we are making progress, there is still a long way to go before AI can truly think like a tactical genius who knows when to hold back and when to strike.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.