Spoiler Alert: Narrative Forecasting as a Metric for Tension in LLM Storytelling
This paper introduces the "100-Endings" metric, which quantifies narrative tension by measuring the frequency of prediction failures in a model's 100 simulated story endings at each sentence, and demonstrates that this approach not only correctly ranks human literary fiction above LLM-generated stories (unlike existing benchmarks) but also guides a structural generation pipeline that significantly improves narrative tension while maintaining EQ-Bench performance.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Problem: AI Writes "Flat" Stories
Imagine you are reading a mystery novel. You turn the page, and your heart races because you don't know if the detective will catch the killer or if the hero will fall off the cliff. That feeling of "What happens next?" is called narrative tension. It's the engine that keeps you reading.
The authors of this paper discovered a weird glitch in Artificial Intelligence (AI). Even the smartest AI models are terrible at writing stories with this kind of tension. They write stories that are grammatically perfect and sound nice, but they feel like watching a movie where the ending is shown in the first five minutes. The AI reveals the secret too early, solves the mystery too quickly, and leaves the reader bored.
Even stranger, when we ask AI to judge stories (acting as a critic), the AI thinks its own flat stories are better than famous human stories from The New Yorker. It's like a robot chef thinking a bowl of plain oatmeal is a Michelin-star meal because it looks neat, while ignoring the fact that it has no flavor.
The Solution: The "100-Endings" Test
To prove that AI stories are boring, the researchers invented a new way to measure tension called the 100-Endings Metric.
Think of a story as a long road trip.
- The Old Way (Rubrics): A critic reads the whole trip and gives it a grade based on how nice the car looked or how polite the driver was.
- The New Way (100-Endings): Imagine you are driving, and at every single mile marker, you stop and ask a psychic: "Based on where we are right now, how will this trip end?"
The psychic tries to guess the destination 100 different times.
- If the story is boring (AI): The psychic guesses the same thing 100 times. "Oh, we're going to the beach." It's obvious. The tension is zero.
- If the story is exciting (Human): The psychic is confused. "Wait, are we going to the beach? Or maybe a mountain? Or did the car break down?" The psychic can't agree on an ending because the story is full of surprises and twists.
The researchers found that human stories (like those from The New Yorker) make the psychic guess wildly different endings, while AI stories make the psychic guess the same ending over and over again.
The Fix: The "Architect" Pipeline
The researchers didn't just want to point out the problem; they wanted to fix it. They realized AI fails because it tries to write the story in one giant leap, like a sprinter who runs the whole race without stopping to think.
So, they built a Story Pipeline that forces the AI to slow down and plan, like a master architect building a house.
- Step 1: The Warm-up (The Tour): The AI reads a famous, high-tension human story (like a classic short story) and studies how the author built the suspense. It learns the "muscle memory" of a good story.
- Step 2: The Blueprint (The Plan): Before writing a single word of the new story, the AI creates a detailed "to-do list." It decides exactly where to hide information, where to drop a bombshell, and how to keep the reader guessing. It's like drawing a map that says, "Here, we will pretend the hero is safe, but actually, he's in danger."
- Step 3: The Build (The Writing): Finally, the AI writes the story, strictly following the blueprint.
The Result:
When they used this "Architect" method, the AI stories became much more like human stories. The "100-Endings" test showed that the AI could finally keep the reader guessing until the very last page. The stories weren't just grammatically correct anymore; they were actually compelling.
The Takeaway
This paper teaches us two big lessons:
- AI is a mimic, not a master: Current AI is great at copying the style of writing (the words), but it doesn't understand the soul of storytelling (the suspense). It solves problems too fast because it's programmed to be helpful, but stories need problems to stay unsolved for a while.
- Structure beats size: You can't just make the AI bigger or smarter to fix this. You have to give it a better structure, like a blueprint. If you force the AI to plan its "surprises" ahead of time, it can write stories that actually make us feel something.
In short: AI can write the words, but it needs a human-like plan to write the magic.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.