The Frequency Confound in Language-Model Surprisal and Metaphor Novelty
This paper demonstrates that lexical frequency, rather than language-model surprisal, is the primary predictor of metaphor novelty, suggesting that previously reported correlations between surprisal and metaphor processing may be confounded by frequency effects.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to understand how humans and computers "think" about language, specifically when it comes to metaphors (like saying "time is money" or "arrested water danced").
Researchers often use a concept called Surprisal to measure how hard it is to understand a word in a sentence. Think of Surprisal like a "shock meter." If you read a sentence and the next word is totally unexpected, the meter spikes high. If the word is obvious, the meter stays low. The theory is that a high shock meter means the word is "novel" (creative) and harder to process.
However, this new paper argues that the "shock meter" might be broken, or at least, it's being tricked by something else: Frequency.
Here is the story of what the researchers found, explained simply:
1. The Confusing Mix-Up
Imagine you are trying to measure how "creative" a new recipe is. You decide to use a "shock meter" to see how surprising the ingredients are.
- The Problem: The researchers found that the shock meter isn't just measuring creativity; it's mostly measuring how rare the ingredients are.
- The Analogy: If you put "salt" in a recipe, it's not surprising because it's common. If you put "dragon fruit," it's surprising because it's rare. But is the recipe creative just because it uses a rare fruit? Not necessarily. The paper suggests that when computers say a metaphor is "novel" (creative), they are often just saying, "Hey, this word is rare!"
2. The "Frequency" vs. "Surprisal" Showdown
The researchers tested this using a massive library of metaphors and 8 different sizes of AI models (from tiny to huge). They compared two things:
- Surprisal: How unexpected the word was in that specific sentence.
- Frequency: How often that word appears in the English language generally.
The Result:
When they asked, "What predicts if a metaphor is considered novel by humans?"
- Frequency was the clear winner. It was a much stronger predictor than Surprisal.
- In fact, the "Surprisal" score was so closely tied to "Frequency" that it was hard to tell them apart. It's like trying to measure the speed of a car, but your speedometer is actually just measuring how much gas is in the tank.
3. The "Baby Model" Mystery
There was a strange pattern the researchers noticed.
- The Expectation: Usually, we think bigger, smarter AI models are better at understanding language.
- The Reality: The smallest AI models (the "baby" models) actually correlated better with human judgments of metaphor novelty.
- The Twist: Why? Because the baby models were "dumber" at predicting words. They were easily shocked by rare words. Since rare words are often the ones humans call "novel," the baby models looked like they were doing a great job.
- The Catch: As the models got bigger and trained longer, they got better at predicting words. They stopped getting "shocked" by rare words. Consequently, their "Surprisal" scores stopped matching human ideas of novelty.
4. The Training Timeline
The researchers watched these AI models as they learned (like watching a student study).
- Early in training: The models were very sensitive to rare words. Their "Surprisal" scores matched human "Novelty" scores perfectly.
- Later in training: As the models learned more, they became less sensitive to rarity. The link between "Surprisal" and "Novelty" broke.
- The Conclusion: The "perfect" match happened early on, but that was only because the model was still very focused on how often words appear (Frequency), not on the deeper context of the sentence.
The Big Takeaway
The paper warns scientists: Don't be fooled.
When we see a small AI model or a model in early training that seems to perfectly predict how creative a metaphor is, it might not be because it understands the meaning or context of the metaphor. It might just be because it is counting how rare the words are.
In short: The "Surprisal" meter is currently too noisy. It's picking up the signal of "Word Rarity" so loudly that we can't hear the signal of "Contextual Surprise." To truly understand how humans process metaphors, we need to separate the idea of "rare words" from "unexpected words in a sentence."
The authors conclude that Lexical Frequency (how common a word is) is the real driver behind why humans rate metaphors as novel, and we need to be careful not to blame "Surprisal" for what is actually just a frequency effect.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.