← Latest papers
💬 NLP

Characterizing Narrative Content in Web-scale LLM Pretraining Data

This paper introduces a novel framework and dataset, NarraDolma, to characterize the fine-grained narrative structure of web-scale LLM pretraining data, revealing that narrative qualities are measurably present but unequally distributed across sources and topics in ways current curation practices overlook.

Original authors: Teagan Johnson, Elliott Ash, Andrew Piper, Maria Antoniak

Published 2026-06-19
📖 5 min read🧠 Deep dive

Original authors: Teagan Johnson, Elliott Ash, Andrew Piper, Maria Antoniak

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are baking a giant cake for a massive party. The recipe calls for "flour," but you don't know what kind of flour you have. Is it fine, silky cake flour? Coarse whole wheat? Is it mixed with sand?

For a long time, scientists building "Large Language Models" (LLMs)—the super-smart AI brains behind tools like chatbots—have been mixing their "flour" (the internet text they train on) without really knowing the specific ingredients. They knew some parts were "toxic" or "about math," but they didn't know much about the stories inside the mix.

This paper is like a team of food scientists who decided to take a microscope to that giant pile of flour. They wanted to answer: How much storytelling is actually in there, and what does it look like?

Here is the breakdown of their investigation:

1. The Recipe Book (The Framework)

The researchers didn't just ask, "Is this a story?" (Yes/No). Instead, they created a detailed flavor profile based on how humans tell stories. They broke "storytelling" down into three main ingredients:

  • The Actor (Agency): Who is doing the acting? Are we seeing the world through their eyes? Do we feel their emotions or hear their thoughts? (Think of this as the "heart" of the story).
  • The Stage (Setting): Where and when is this happening? Is it a vague "somewhere," or is it a specific, dusty room in 1920s Paris? (Think of this as the "backdrop").
  • The Plot (Events): What actually happens? Do things change? Is there a cause-and-effect chain, or just a list of things? (Think of this as the "action").

They turned these into 11 specific "dials" or sliders that can be rated from 1 to 5.

2. The Taste Test (The Data)

They took a massive library of internet text called DOLMA (which is huge—3 trillion words, like a library that would take you a lifetime to read).

  • Step 1: They couldn't read it all. So, they built a robot assistant (a model called NARRABERT) to help them.
  • Step 2: First, they hired human experts to read 400 random snippets and rate them on those 11 dials. This was the "gold standard."
  • Step 3: They taught the robot assistant to mimic the humans.
  • Step 4: The robot went through 3 million snippets and rated them all. They created a new map called NARRADOLMA.

3. The Surprising Findings

When they looked at the map, they found three big things:

A. Stories aren't just "on" or "off."
It's not like a light switch where a text is either a story or not. It's more like a dimmer switch. Some texts are very dim (boring facts), some are bright (dramatic novels), and most are somewhere in between. The researchers found that "storytelling" in the internet is a continuous, multi-dimensional landscape, not a simple category.

B. The "Flour" is unevenly mixed.
If you think "Reddit" is full of stories and "Wikipedia" is full of facts, you're mostly right, but it's more complex.

  • Reddit and Books are like a spice rack full of "inner feelings" and "character thoughts." They are high on the "Actor" dials.
  • Wikipedia and News are like a map. They are high on "Events" and "Time/Place" but low on "Feelings."
  • Travel and Food blogs are like a painting. They are high on "Sensory details" (smells, sights) but low on "Conflict."

C. You can't just "add more stories" by adding more books.
The researchers found that even within a single source (like Reddit), the "storytelling" varies wildly. Some Reddit posts are pure drama; others are just technical questions.

  • The Metaphor: Imagine you want to make your cake more "fluffy." You might think, "I'll just add more cake flour!" But if that bag of flour actually contains a mix of cake flour, sand, and gravel, just adding more of the bag doesn't guarantee a fluffier cake. You need to know which specific part of the bag is the good flour.

4. Why This Matters

The paper concludes that the people who build AI models have been treating their data sources (like "Reddit" or "News") as if they are all the same kind of story-teller. They aren't.

If an AI is trained mostly on "News" (high events, low feelings), it might get good at reporting facts but bad at understanding human emotions. If it's trained mostly on "Reddit" (high feelings, low structure), it might get good at empathy but bad at logical cause-and-effect.

The Bottom Line:
This paper didn't invent a new AI or fix a broken one. Instead, it built a measuring tape for the internet's stories. It showed us that the "story" part of our data is messy, varied, and unevenly distributed. Before we can teach AI to tell better stories, we first have to understand exactly what kind of stories are currently in the mix.

The researchers have released their measuring tape (the model) and their map (the dataset) for anyone else to use, so we can all start baking better cakes in the future.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →