Computational Representations of Character Significance in Novels
This paper proposes a novel six-component structural model of character that incorporates discussion by other characters, demonstrating through computational analysis of 19th-century British realist novels that this approach offers new insights into literary theories of character centrality and gendered dynamics.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a novel not just as a story, but as a bustling city. For a long time, literary critics and computer scientists tried to figure out who the "important" people in that city were by simply counting how many times they walked down the street (appeared in a scene) or how many times they shook hands with others (had a conversation).
This paper argues that this method is like judging a person's importance in a city only by how often they are seen in public. It misses the fact that someone might be the most talked-about person in town, even if they stay home all day.
Here is a breakdown of what the researchers did, using simple analogies:
1. The New "Six-Point" Scorecard
Instead of just counting "street appearances," the team (a mix of computer scientists and literature experts) created a new six-part scorecard to measure a character's "space" in the story. Think of it like a report card for a character's presence:
- Name (N): How many times is their name written down? (Like seeing their name on a mailbox).
- Communication (C): How much do they speak or write? (Like the volume of their voice in a town square).
- Interiority (I): How much do we get inside their head to hear their thoughts and feelings? (Like reading their private diary).
- Action (A): What physical things do they do? (Like building a house or running a race).
- Discussion by Others (DC): How much do other characters talk about them? (This is the new big one. It's like how often people at a party whisper about someone who isn't even there).
- Narrator Description (DN): How much does the storyteller describe them? (Like a narrator pausing to say, "Look at his sad eyes").
2. The "Gossip Map" (The New Network)
The researchers built a special kind of map (a graph) based on the "Discussion" part of the scorecard.
- Old Maps: Previous maps showed who stood next to whom (co-occurrence) or who talked to whom (dialogue). These are like two-way streets.
- The New Map: This map shows who is talking about whom. This is a one-way street.
- Analogy: Imagine a town where Alice never speaks, but everyone else is constantly talking about her. In an old map, she might look invisible. In this new "Gossip Map," she is the center of the universe because everyone is pointing at her.
3. The Experiment: Teaching Computers to Read Like Critics
The team tested two ways to get computers to fill out this six-part scorecard for 19th-century British novels (like Pride and Prejudice and Jane Eyre):
- Method A (The "BookNLP" Pipeline): A specialized, rule-based tool designed specifically for literature. It's like a seasoned librarian who knows exactly where to look for specific types of sentences.
- Method B (The "LLM" Approach): Using powerful AI chatbots (like GPT-4) to read the text and guess the counts.
The Result: The specialized librarian tool (BookNLP) was much more accurate and consistent than the AI chatbot. The chatbot often got confused, especially with minor characters or when the story was told from a first-person perspective (where the narrator is also a character). The researchers decided to use the "librarian" tool for their big study.
4. What They Discovered
Using this new scorecard and the "Gossip Map," they analyzed 64 classic novels and found some surprising things:
- The "One vs. The Many" Theory: A famous theory suggests novels are built around one main hero and a crowd of minor characters fighting for space. The researchers found this is mostly true, but the "One" is often actually a small cluster of 2 or 3 people who share the spotlight, rather than just a single hero.
- Being Talked About vs. Talking: In Pride and Prejudice, Mr. Darcy is the most "talked about" character (he has the highest "In-Degree" on the gossip map). However, Elizabeth Bennet is the one doing the most talking about others (highest "Out-Degree"). This shows that being the topic of conversation is different from being the speaker.
- Gender and Gossip: They looked at who talks about whom based on gender. They found that women in these novels talk about other women more than men talk about other men. Also, women were more likely to be the ones discussing others, while men were often the ones being discussed. This revealed a hidden social layer that the old "who stood next to whom" maps completely missed.
Summary
The paper introduces a new way to measure character importance that goes beyond just "who shows up." By counting how much characters think, act, and—crucially—how much they are talked about by others, the researchers created a more nuanced map of a novel's social world. They proved that computers can do this, but they need the right tools (specialized software rather than just general AI) to get the details right. This allows us to see that in these classic stories, the "main character" is often a complex web of relationships, not just a single hero.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.