← Latest papers
💬 NLP

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives

This paper introduces NCP-Bench, a benchmark of 100 narrative environments designed to evaluate Long-Horizon Consistency in interactive storytelling, revealing that even state-of-the-art LLMs struggle significantly with maintaining logical commitment preservation and narrative integrity against unconstrained user interventions.

Original authors: Yingpeng Ma, Jianhao Yan, Bei Shi, Ka Hou Kam, Runnan Wang, Xuebo Liu, Yulong Chen, Yue Zhang, Derek F. Wong

Published 2026-08-11
📖 4 min read☕ Coffee break read

Original authors: Yingpeng Ma, Jianhao Yan, Bei Shi, Ka Hou Kam, Runnan Wang, Xuebo Liu, Yulong Chen, Yue Zhang, Derek F. Wong

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are playing a massive, endless role-playing game where the story never ends, and you can do absolutely anything you want. In this world, you might decide to punch a dragon, steal a king's crown, or convince a villain to become a baker. The magic that makes this possible is a type of super-smart computer brain called a Large Language Model (LLM). Think of these models as incredibly talented improvisational actors who have read almost every book ever written. They are great at sounding natural, creating atmosphere, and keeping the conversation flowing. But there's a catch: while they are excellent at sounding like they know the story, they often struggle to actually remember the rules of the world they are building. If the computer says "the door is locked," a few minutes later it might let you walk right through it as if it were open, or it might forget that you already found the key. This is a big problem for game designers and storytellers who want a world that feels real and consistent, no matter how wild the player gets.

This paper, titled "Can LLM Agents Stick to the Script?", dives deep into this exact problem. The researchers created a special testing ground called NCP-Bench (Narrative Commitment Preservation Benchmark) to see if these AI storytellers can keep their promises over a long, messy game session. They set up 100 different story worlds based on famous movie plots, like Iron Man or The Bourne Identity. In each world, the AI acts as the "Game Master," responsible for running the show, while a second AI acts as a "troublemaker" player, trying to skip ahead, break the rules, or force the story to go in impossible directions. The goal was to see if the Game Master could stay consistent—keeping track of facts, remembering what happened, and sticking to the plot's required milestones—even when the player was trying to break the game.

The results were a bit of a reality check for the AI community. The study found that even the most advanced, state-of-the-art AI models are surprisingly bad at sticking to the script when things get complicated. While these models can write beautiful, fluent, and exciting stories, they often lose the plot when the game goes on for too long. In the tests, the best-performing model (a version of GPT-5.2) managed to survive without making a logical error for only 42% of the time after just 20 turns of interaction. As the game went on, the survival rate dropped sharply. By the time the game hit the 100-turn limit, almost no model could finish the whole story without contradicting itself or forgetting a major plot point.

The researchers discovered that the most common mistake wasn't that the AI was being mean or ignoring the player; it was that the AI simply forgot the facts it had established earlier. About 40% to 68% of the failures were "fact conflicts," where the AI would say something that directly contradicted what it said ten minutes ago. For example, if the story established that a character was injured, the AI might later describe them running a marathon without a scratch. Even when the AI tried to be helpful and let the player skip boring parts of the story, it often did so by rewriting history in a way that broke the logic of the world.

Interestingly, the paper also tested a special kind of AI designed with a "memory" system to help it remember things better. While this memory system helped the AI last a little longer in the game, it didn't solve the problem. In fact, the memory system sometimes made the AI worse at listening to the player's specific actions, because it was too busy summarizing the story into big chunks and missing the small details. The study suggests that simply making AI smarter or giving it more memory isn't enough; we need a new way to teach these models how to treat the story's rules as unbreakable laws, not just suggestions.

In short, the paper concludes that while our current AI storytellers are fantastic at improvising and sounding cool, they are currently terrible at playing the long game. They can't reliably stick to the script when the player tries to break the rules. The authors hope that by providing this benchmark, other researchers can build better tools to fix this, so that one day, we can have truly endless, consistent, and magical interactive stories where the world makes sense no matter what we do.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →