← Latest papers
💬 NLP

Surprisal and Metaphor Novelty Judgments: Moderate Correlations and Divergent Scaling Effects Revealed by Corpus-Based and Synthetic Datasets

This study demonstrates that while language model surprisal shows moderate correlations with metaphor novelty annotations, it exhibits divergent scaling behaviors—decreasing in predictive power with larger models on corpus-based data but increasing on synthetic data—highlighting its limitations as a comprehensive metric for linguistic creativity.

Original authors: Omar Momen, Emilie Sitter, Berenike Herrmann, Sina Zarrieß

Published 2026-01-27
📖 4 min read☕ Coffee break read

Original authors: Omar Momen, Emilie Sitter, Berenike Herrmann, Sina Zarrieß

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to guess the next word in a sentence. If the sentence is "The cat sat on the...", your brain (and a computer program) can easily guess "mat." That was easy to predict, so it wasn't very surprising. But if the sentence is "The cat sat on the... jellyfish," your brain does a double-take. That word was hard to predict, so it has high "surprisal."

This paper asks a simple question: Can this feeling of "surprise" help us tell the difference between a boring, old metaphor and a fresh, creative one?

The Big Idea: Predictability vs. Creativity

Metaphors are like mental bridges. Sometimes the bridge is a well-worn path everyone knows, like saying "I attacked his argument." Everyone knows you don't use a weapon on a speech; it's a standard phrase. This is a conventional metaphor.

Other times, the bridge is brand new and shaky, like saying "The arrested water." You have to stop and think: How can water be arrested? It's not in the dictionary. This is a novel metaphor.

The researchers wanted to see if computer programs (called Language Models) could use their "surprise meter" to spot these new, creative bridges. The theory was: If a word is hard to predict, maybe it's a creative metaphor.

The Experiment: Two Different Worlds

The team tested this idea using two very different types of data, and the results were like a tale of two cities:

1. The "Real World" City (Corpus-Based Data)
They looked at real sentences taken from books, news, and conversations.

  • The Result: Here, the "surprise meter" worked okay, but with a twist. Smaller, simpler computer models were actually better at spotting the creative metaphors than the giant, super-smart models.
  • The Analogy: Imagine a small, local detective who knows the neighborhood well. They notice the odd, new things because they aren't overthinking it. The giant detective (the big AI) has seen so much data that it starts expecting weird things everywhere, making it less sensitive to what's actually "new" in a specific sentence. The bigger the model got, the worse it became at this specific task.

2. The "Toy Factory" City (Synthetic Data)
They also used sentences that were carefully built in a lab (or by another AI) where the only difference was whether the metaphor was old or new.

  • The Result: Here, the rules flipped. The bigger, smarter models got much better at spotting the creative metaphors.
  • The Analogy: In a controlled toy factory, the giant detective shines. Because the toys are perfect and the rules are strict, the massive brain can use its huge power to see the subtle differences that the small detective misses.

The "Cloze" Trick: Looking Both Ways

The researchers also tried a new trick. Standard "surprise" only looks at the words before the target word (like reading a sentence from left to right). But to understand a metaphor, you often need to see what comes after it too.

They created a "Cloze" method (like a fill-in-the-blank game). They hid the metaphorical word, showed the model the whole sentence (including the end), and asked, "What word goes here?"

  • The Result: This "looking both ways" approach generally made the computer smarter at spotting the metaphors. It was like giving the detective a map of the whole street instead of just the corner they were standing on. However, this trick worked best for the "Real World" sentences and sometimes confused the models on the "Toy Factory" sentences.

What About Teaching the AI to Talk Like a Human?

The team also tried using "Instruction-Tuned" models (AI that has been taught to follow human instructions and chat nicely).

  • The Result: Surprisingly, making the AI talk more like a human did not make it better at spotting creative metaphors. In fact, for some models, it made them worse at this specific math problem. Being "chatty" didn't help them be "creative."

The Bottom Line

The paper concludes that "surprisal" (how unexpected a word is) is a decent tool for finding creative metaphors, but it's not a magic wand.

  • It works moderately well (about 40-50% correlation), meaning it's helpful but far from perfect.
  • Context matters: Whether the AI is looking at messy real-world text or clean, lab-made text changes everything.
  • Size matters: Bigger isn't always better; sometimes smaller is sharper, depending on the data.

In short, while computers can get a hint of human creativity by measuring how "surprised" they are, they still struggle to fully grasp the magic of a truly novel metaphor. They need more than just a surprise meter to understand the art of language.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →