← Latest papers
💬 NLP

Attention Flows: Tracing LLM Conceptual Engagement via Story Summaries

This paper evaluates the ability of state-of-the-art LLMs to comprehend long-form narratives by comparing their novel summaries against human-written ones, revealing that models struggle with holistic integration and disproportionately focus on the ends of texts, a finding linked to their attention mechanisms.

Original authors: Rebecca M. M. Hicke, Sil Hamilton, David Mimno, Ross Deans Kristensen-McLachlan

Published 2026-04-09
📖 5 min read🧠 Deep dive

Original authors: Rebecca M. M. Hicke, Sil Hamilton, David Mimno, Ross Deans Kristensen-McLachlan

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: Can AI Actually "Read" a Whole Book?

Imagine you give a super-smart robot a 500-page novel and ask it to write a summary. In the past, robots were like people with very short attention spans; they could only remember the first few pages and the very last page. If you asked them about the middle, they'd just guess.

Recently, these robots (Large Language Models, or LLMs) have gotten "memory upgrades." They can now hold millions of words in their "mind" at once. The big question this paper asks is: Just because they can hold the whole book, does that mean they actually understand it?

The authors argue that "retrieving" a fact (finding a needle in a haystack) is easy, but "comprehending" a story (understanding the whole plot) is hard. They call the failure to understand the middle of a story "Context Rot." It's like reading a book but your brain starts to rot or decay in the middle chapters, leaving you with a clear memory of the beginning and end, but a fuzzy mess in the middle.

The Experiment: The "Story Map" Test

To test if these robots truly understand stories, the researchers didn't just ask for a summary. They did something clever: They created a "Story Map."

  1. The Human Baseline: They took 150 classic novels (like Dracula or Pride and Prejudice) and looked at the summaries humans wrote on Wikipedia.
  2. The Robot Summaries: They asked 9 different AI models to write summaries of the same books.
  3. The Alignment: This is the tricky part. They took every single sentence from the summaries and asked: "Which specific chapter of the original book does this sentence describe?"

Think of it like this: If a human reads a book, they create a mental map where they visit every room (chapter) to see what happened. The researchers wanted to see if the AI's map looked the same as the human's map.

What They Found: The "End-Weighting" Problem

The results were revealing. While the AI summaries sounded okay, their "mental maps" were very different from humans.

  • Humans are Tour Guides: When humans summarize a story, they tend to visit every room in the house. They spend a little time in the beginning, a lot in the middle, and a bit at the end. They cover the whole journey.
  • AI are "Recency Bias" Fans: The AI models acted like they were only interested in the beginning and the end of the book. They largely ignored the middle.
    • The Metaphor: Imagine you are describing a road trip to a friend. A human would say, "We drove through the mountains, got lost in the desert, saw a cool canyon, and then arrived at the beach." The AI would say, "We started at the mountains... and then we were at the beach!" It skipped the entire middle of the trip.

The paper calls this "End-Weighting." The AI models seem to think the most important parts of a story are the very first thing that happens and the very last thing that happens, treating the middle chapters as boring filler.

Other Weird Differences

The study also found some other quirks in how AI writes compared to humans:

  • The "High-Level" Summary: AI summaries were often shorter and more vague. Instead of saying "The hero fought the dragon in the cave," the AI might say "The hero faced a great challenge." It's like a robot giving you the "Cliff's Notes" version without the details.
  • Less Linear: Humans usually tell stories in order (Chapter 1, then 2, then 3). AI summaries often jumped around, mixing events from different times, as if they were looking at the whole book at once rather than reading it page-by-page.
  • The "Lost in the Middle" Attention: The researchers looked inside one of the AI models (Qwen 3.5) to see how its "attention" worked. They found that the model literally paid less "attention" to the middle chapters. It was like a flashlight that was bright at the start and end of the book but dim in the middle.

Why Does This Matter?

This is a big deal because we are starting to use AI to analyze huge amounts of text—legal documents, medical records, and history books.

If an AI is good at finding a specific date in a document (the "Needle in a Haystack"), but bad at understanding the story of the document (the "Context Rot"), then we can't fully trust it to make complex decisions. It might miss the crucial plot twist that happens in Chapter 40 of a 500-page contract.

The Takeaway

The paper concludes that while AI has gotten much better at holding long texts in its "memory," it hasn't quite learned how to sustain attention across the whole story. It's like a student who can memorize the first and last page of a textbook perfectly but falls asleep during the middle 400 pages.

The researchers released their data (the 150 books and all the summaries) so other scientists can try to fix this "Context Rot" and teach AI to pay attention to the whole story, not just the beginning and the end.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →