← Latest papers
💬 NLP

Computational Analysis of Semantic Connections Between Herman Melville Reading and Writing

This study employs computational semantic similarity analysis using BERTScore to investigate the influence of Herman Melville's reading on his writing, demonstrating that the method effectively identifies known literary connections and highlights new passages for qualitative examination.

Original authors: Nudrat Habib, Elisa Barney Smith, Steven Olsen Smith

Published 2026-03-17
📖 5 min read🧠 Deep dive

Original authors: Nudrat Habib, Elisa Barney Smith, Steven Olsen Smith

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to solve a mystery: Did Herman Melville, the famous author of Moby Dick, secretly copy ideas from books he read, or did he just happen to think of the same things?

For a long time, literary detectives (scholars) had to read thousands of pages of Melville's work and hundreds of books from his personal library, manually hunting for clues. It was like trying to find a specific needle in a haystack by looking at every single piece of hay with a magnifying glass.

This paper introduces a new, high-tech detective tool to help solve the case faster. Here is the story of how they did it, explained simply.

The Big Idea: A "Semantic" Magnifying Glass

The researchers wanted to see if Melville's writing "smelled" like the books he owned. But they didn't just look for exact word-for-word copies (like a plagiarism detector). They wanted to find ideas that were similar, even if Melville changed the words.

Think of it like this: If you read a recipe for "Chocolate Cake" and then write a story about "a delicious, sweet, brown dessert," a simple word-search might miss the connection. But a smart detective knows that "Chocolate Cake" and "delicious, sweet, brown dessert" are essentially the same idea.

The tool they used is called BERTScore. Imagine it as a super-smart robot that reads two texts and asks, "Do these two sentences mean the same thing, even if they use different words?"

How They Set Up the Experiment

  1. The Suspects: They picked four famous pairs of books. On one side, they had a book Melville wrote (like Moby Dick). On the other, they had a book Melville definitely read (like a book about whales).
  2. The "Ground Truth": Before they turned on the robot, human experts had already pointed out specific spots where Melville definitely borrowed ideas. These were the "known clues."
  3. The Two Ways to Look: To make sure they didn't miss anything, they looked at the text in two different ways:
    • The "Phrase" Lens (5-grams): They broke the text into tiny chunks of five words. This is like looking for specific phrases or catchy sayings Melville might have stolen.
    • The "Sentence" Lens: They looked at whole sentences. This is like looking at the bigger picture of an entire thought.

What They Found

The robot did a great job, but it also taught them some interesting lessons about how writing works.

1. The Robot Found the Known Clues
When the experts said, "Hey, Melville copied this part of Typee from this book about the South Seas," the robot agreed. It gave them high scores, confirming that the tool works. It could see that even though Melville changed some words, the meaning was still a perfect match.

2. The Robot Found New Clues (The "Eureka" Moments)
The robot didn't just stop at the known clues. It scanned the whole text and found new places where Melville's writing looked suspiciously like his reading. These were spots the human experts hadn't noticed yet. This is like the detective finding a hidden fingerprint on a window that no one else saw.

3. The "One-to-Many" Puzzle
Here is where it got tricky. Sometimes, Melville would read one long, complicated sentence in a source book and turn it into three shorter sentences in his own book.

  • The Problem: When the robot compared the source sentence to just one of Melville's new sentences, the score was low. It was like comparing a whole pizza to just one slice; they don't look the same size, even though the slice came from the pizza.
  • The Lesson: The robot is great at spotting direct matches, but it needs human help to realize that sometimes an author spreads one big idea across many small sentences.

4. The "Generic" Trap (False Alarms)
The robot sometimes got excited about boring sentences. For example, if a source book said, "The man walked into the room," and Melville wrote, "The knight visited his workers," the robot gave them a high score.

  • Why? Because "walking," "visiting," "men," and "workers" are common words that appear everywhere. The robot thought, "Hey, these are similar!"
  • The Reality: A human knows these are just generic phrases that don't prove Melville copied the specific story. The robot needs a human to say, "Okay, that's just a coincidence, not a clue."

The Takeaway

This study is like building a new kind of metal detector for literature.

  • Before: Scholars had to walk the beach with their eyes, hoping to spot a coin.
  • Now: They have a metal detector that beeps when it finds buried treasure (similar ideas).

The detector isn't perfect—it sometimes beeps at a soda can (generic words) or misses a coin if it's buried under a pile of sand (complex rewriting). But it is incredibly fast and helps scholars find the real treasures much quicker.

In short: By using AI to compare what Melville read with what he wrote, the researchers proved that computers can help us understand how authors borrow, twist, and improve upon the stories of others. It doesn't replace the human expert; it just gives them a powerful flashlight to see in the dark.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →