← Latest papers
💬 NLP

Generative AI & Fictionality: How Novels Power Large Language Models

This paper argues that novels significantly shape the outputs of generative AI by providing rich linguistic patterns that models leverage to simulate social and communicative phenomena, necessitating that cultural analysis now account for the influence of computational training data.

Original authors: Edwin Roland, Richard Jean So

Published 2026-03-03
📖 5 min read🧠 Deep dive

Original authors: Edwin Roland, Richard Jean So

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: Who is Teaching the Robot to Talk?

Imagine you are trying to teach a robot how to be a human. You have two choices for its schoolbooks:

  1. The Encyclopedia: A massive library of facts, news, and Wikipedia articles.
  2. The Novel Library: A massive collection of fiction, romance, sci-fi, and mystery books.

Most people assume that to make a smart AI, you need the Encyclopedia. But the authors of this paper, Edwin Roland and Richard Jean So, discovered something surprising: The Novel Library is actually the secret sauce.

They found that modern AI (like ChatGPT) isn't just learning facts; it's learning how to be a person because it has read millions of novels.


The Experiment: The "Wikipedia Robot" vs. The "Novel Robot"

To prove this, the researchers built two digital brains (called BERT models) and gave them different diets:

  • The Wiki Robot: Fed only Wikipedia articles (facts, dates, policies).
  • The Full Robot: Fed Wikipedia plus a huge collection of novels.

Then, they asked both robots to finish sentences and generate stories. Here is what happened:

1. The "Gentleman Scientist" vs. The "Dramatic Character"

  • The Wiki Robot sounds like a boring, polite encyclopedia entry. It talks about elections, military tech, and bird nesting habits. It's accurate, but it has no soul. It's like a librarian reading a list of ingredients.
  • The Full Robot sounds like a character in a movie. It uses words like "you," "I," "want," and "could." It talks about feelings, relationships, and what people are thinking. It's like a novelist writing a scene.

The Takeaway: Reading fiction taught the AI how to simulate a human mind, not just a database.

2. The "Ghost" in the Machine

The researchers noticed something weird about names.

  • If you ask the AI about a specific person (like "Michelle" or "Ronnie"), it actually gets worse at predicting what comes next.
  • But if you ask it about pronouns (like "he," "she," "they," "you"), it gets better.

The Analogy: Think of a specific name as a real person standing outside the room. The AI doesn't know them. But a pronoun is like a ghost inside the story. The AI learned from novels that pronouns are the glue that holds a character's identity together within the story. It learned that "he" isn't just a word; it's a placeholder for a person with a history, a body, and a mind.

3. The "Drama Detector"

When the researchers looked at the best parts of the novels the AI learned from, they found a pattern. The AI didn't just learn random words; it learned high-stakes moments.

The AI learned that the most important parts of a story are:

  • Two people arguing or falling in love.
  • A character realizing something about themselves (a "clenching of the jaw").
  • A moment where the future changes (a climax).

The Metaphor: Imagine the AI is a student who only studied the most dramatic scenes of a play. It didn't learn the boring parts where people just sit and eat lunch. It learned that human life is defined by conflict, intention, and emotional turning points.


Why Does This Matter?

The paper argues that we are entering a new era where AI is the new author.

  • The Old Way: A human writer sits down, thinks about their intentions, and writes a story. We analyze their mind to understand the book.
  • The New Way: An AI scans millions of books, mixes them all together, and spits out new text. It doesn't have a "mind" or "intentions" in the human sense. It's a "stochastic parrot" (a fancy way of saying it repeats patterns it heard).

The Problem:
If 90% of the internet content in the future is written by AI, and that AI learned its "personality" from a specific mix of novels, then our culture is being shaped by the books the AI read.

If the AI learned mostly from romance novels, it might think all human relationships are dramatic and full of misunderstandings. If it learned mostly from thrillers, it might think the world is full of danger.

The Final Lesson: Check the Recipe

The authors conclude that we can't just look at the final dish (the AI's output) to understand what's happening. We have to look at the recipe (the training data).

  • Analogy: If you eat a cake and it tastes like strawberries, you can't just say "It's a strawberry cake." You have to ask, "Did the baker put strawberries in, or did they just use a strawberry-flavored powder?"
  • The Point: To understand our future culture, we need to audit the "ingredients" (the data) that the AI is eating. We need to know which novels, which news sites, and which Reddit threads are teaching the robot how to be human.

In short: Fiction didn't just entertain us; it accidentally taught the robots how to be people. And now, those robots are writing the stories of our future.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →