Clozing the Gap: Exploring Why Language Model Surprisal Outperforms Cloze Surprisal
This paper argues that language model surprisal outperforms cloze surprisal in predicting processing effort due to its superior resolution, ability to distinguish semantically similar words, and accurate handling of low-frequency words, while calling for improved cloze methodologies and further research into human prediction sensitivity.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to guess the next word in a sentence. Sometimes the answer is obvious, and sometimes it's a total surprise. Psychologists have long used a game called the "Cloze Task" to measure how predictable a word is. In this game, you show a person a sentence with the last word missing (e.g., "The cat sat on the...") and ask them to fill in the blank. If most people say "mat," the word is highly predictable. If they say many different things, it's unpredictable.
For decades, this human game was the gold standard for understanding how our brains process language. But recently, a new contender has entered the ring: Language Models (LMs), like the AI behind this very explanation. These computers are trained on billions of sentences and can calculate the exact mathematical probability of what word comes next.
The Big Discovery:
The paper finds that when predicting how fast humans read a word, the computer's math (LM probabilities) is a much better guess than the human game (Cloze Task). The computer's predictions line up perfectly with how long our eyes pause on a word, while the human game is a bit clunky and misses the mark.
But the authors ask a crucial question: Is the computer winning because it's smarter, or just because it's playing a different game?
To find out, the researchers ran three "experiments" where they tried to make the computer play by the human's rules. They tested three reasons why the computer might be winning:
1. The "Pixel Resolution" Test (Hypothesis 1)
The Metaphor: Imagine the human game is a low-resolution, pixelated photo. You can see the general shape of a face, but you can't see the freckles. The computer, however, is a 4K ultra-HD photo with every tiny detail visible.
The Experiment: The researchers forced the computer to only look at a limited number of samples, just like the human game (which usually only asks about 40–90 people). They "pixelated" the computer's high-definition math to match the low-resolution human data.
The Result: When the computer's vision was blurred to match the human game, it stopped being the better predictor. This suggests the computer wins partly because it has a much higher "resolution"—it can distinguish between tiny nuances that the human game misses because it doesn't have enough people to ask.
2. The "Twin Brothers" Test (Hypothesis 2)
The Metaphor: Imagine the sentence is "He sat on the..." and the possible answers are "couch" and "sofa." To a human, these are twins; they mean the same thing. If a human guesses "couch," they probably also thought of "sofa." But the computer treats them as two completely different strangers with different scores.
The Experiment: The researchers told the computer to group similar words together. If "couch" and "sofa" are in the same group, the computer had to give them the same score.
The Result: When the computer was forced to treat similar words as a single team, its predictive power dropped. This means the computer's advantage comes from its ability to make extremely fine-grained distinctions between words that humans might see as the same.
3. The "Rare Word" Test (Hypothesis 3)
The Metaphor: Imagine a sentence where the next word is "flabbergasted." In a human game, nobody might guess that rare word because they've never heard it often. But the computer, having read the entire internet, knows exactly how likely that rare word is.
The Experiment: The researchers told the computer to ignore all rare words and only guess common ones, just like a human participant who wouldn't guess a word they've never seen.
The Result: When the computer was forced to ignore rare words, it got worse at predicting reading times. This shows the computer is better at handling the "long tail" of rare words that humans simply don't produce in the game.
The Conclusion: A Call for Better Tools
The paper concludes that the computer isn't necessarily "thinking" like a human; it's just a better measuring stick because it has higher resolution, can spot subtle differences, and handles rare words well.
However, the authors warn us not to just throw away the human game. They argue that the human game is "low resolution" because it's too expensive and slow to get thousands of answers for every sentence.
The Takeaway:
We need to upgrade our human experiments. Instead of just asking people to guess the next word once, we need new ways to measure human prediction that are as sensitive and detailed as the computer's math. Until we do that, we can't be 100% sure if the computer is predicting how humans think, or if it's just a really good calculator that happens to match our reading speeds for the wrong reasons.
In short: The computer is currently the better ruler for measuring reading speed, but it might be measuring things humans don't actually notice. We need to build a better ruler for humans to see if they really see the world the same way the computer does.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.