Why are language models less surprised than humans? Testing the Parse Multiplicity Mismatch Hypothesis
This paper tests the hypothesis that language models' underestimation of human processing difficulty in syntactic ambiguity arises from their ability to maintain more simultaneous sentence parses than humans, but finds that while reducing parse multiplicity increases predicted garden-path effects, it fails to fully account for the magnitude of human reading time differences.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are reading a story, and suddenly you hit a sentence that tricks your brain. This is called a "garden path" sentence.
Here is an example: "The girl fed the lamb was upset because she asked for beef."
When you first read "The girl fed the lamb," your brain naturally assumes the girl is the one doing the feeding. It's the most common way to understand it. But then you hit the word "was." Suddenly, your brain realizes it was wrong! The sentence actually means the girl was fed (someone fed her), and she was upset about the beef. That moment of confusion and re-reading takes time. Psychologists call this a "garden path effect."
The Big Question
Scientists have built powerful computer programs called "Language Models" (like the ones that write this text) that are very good at guessing the next word in a sentence. A theory called "Surprisal Theory" suggests that the harder a word is to predict, the longer it takes humans to read it.
The problem? When scientists tested these computers on garden path sentences, the computers were too calm. They didn't seem surprised enough by the trick. They predicted the confusion would be tiny, but humans actually experience a massive mental "stumble."
The Hypothesis: The "Flashlight" Theory
The authors of this paper wondered: Why are the computers so calm?
They guessed it might be because computers are too smart. When a computer reads "The girl fed the lamb," it might be holding dozens of different interpretations in its head at the same time. It's thinking: "Maybe she's feeding the lamb. Maybe the lamb is feeding her. Maybe it's a different kind of lamb." Because it's looking at all these possibilities at once, the word "was" doesn't feel like a shock; it just feels like one of the many options it was already considering.
Humans, on the other hand, might be like a flashlight in a dark room. We can only shine our light on one or two paths at a time. We commit to the "girl is feeding" idea. When we hit "was," our flashlight is pointing the wrong way, and we have to spin around and find the new path. That spinning takes time and effort.
The authors called this the "Parse Multiplicity Mismatch Hypothesis." In simple terms: Computers are less surprised because they are looking at too many possibilities at once, while humans are forced to pick just one.
The Experiment: Turning Down the Flashlight
To test this, the researchers used a special type of computer model that can be told exactly how many "paths" to look at. They ran the models through the garden path sentences with different settings:
- High Multiplicity: The computer looks at 1,000 different interpretations at once (like a super-bright, wide flashlight).
- Low Multiplicity: The computer is forced to look at only 1 or 2 interpretations (like a dim, narrow flashlight).
The Results
- They were partially right: When they forced the computer to look at fewer paths (narrower flashlight), it did get more surprised. The "stumble" in the computer's prediction got bigger.
- But they were mostly wrong: Even when they forced the computer to look at only one path (the wrong one, just like a human), the computer's surprise was still tiny compared to the human reaction.
The computer's "stumble" was still about 40 times smaller than the human stumble.
The Conclusion
The paper concludes that simply making computers "dumber" or forcing them to look at fewer options doesn't fix the problem. The gap between how computers and humans process language is too big to be explained just by how many ideas they can hold in their heads at once.
The authors suggest that human brains might have an extra step that computers don't have: a specific "re-analysis" mechanism. When humans get stuck, we don't just feel a little surprised; we actively tear down our old understanding and rebuild a new one. The current computer models, even when limited, just don't seem to have that "tear down and rebuild" button. They just keep guessing, and they guess too easily.
In a Nutshell
The researchers tried to see if computers were less surprised by tricky sentences because they were "thinking too much" (holding too many ideas). They found that even when they forced the computers to "think less," the computers were still way too calm compared to humans. This means the difference between human and computer reading isn't just about how many ideas we can juggle; it's about something deeper in how we handle being wrong.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.