To model human linguistic prediction, make LLMs less superhuman
The paper argues that large language models have become less effective at modeling human reading behavior because their superior predictive accuracy, driven by extensive training data and memory, exceeds human capabilities, and advocates for developing models with human-like memory constraints to better align with human linguistic prediction.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are reading a mystery novel. As you turn the page, your brain is constantly guessing what word comes next. If the story says, "The detective found a hidden key in the old _______," you might guess "box," "drawer," or "safe." These guesses aren't just random; they happen so fast that they actually change how quickly you read the sentence. If your guess is right, you breeze through; if it's wrong, you stumble.
For a long time, scientists thought Large Language Models (LLMs)—the super-smart AI behind chatbots—were perfect partners for studying how humans read. After all, these AIs also guess the next word. But here is the twist: The smarter the AI gets at guessing, the worse it becomes at mimicking how a real human reads.
This paper argues that current AI models have become "superhuman" in a way that actually breaks the comparison. They are too good, and their "superpowers" are two types of memory that humans simply don't have.
The Two Superpowers That Make AI "Too Good"
1. The Photographic Encyclopedia (Long-Term Memory)
Imagine you are taking a test on history. You might know that Elvis Presley was born in Tupelo, Mississippi, because you read it in a book once. But maybe you forgot it, or maybe you never learned it at all.
Now, imagine a student who has read every single book, website, and article ever written, and has memorized every fact in them perfectly. That is the AI.
- The Problem: When the sentence says, "Elvis was born in the city of _______," the AI instantly knows the answer is "Tupelo" with 100% certainty. A human reader, however, might not know the answer at all, or might guess "Memphis" or "Nashville" based on what they do know.
- The Result: Because the AI knows facts humans don't, its "guess" is too sharp and too specific. It doesn't struggle with the sentence the way a human does. To be a good model of human reading, the AI needs to be "dumber" about facts, forgetting things just like we do.
2. The Perfect Replay Button (Short-Term Memory)
Now, imagine you are reading a long story with many characters. You read a sentence about a character named "Barnaby" on page 1. You read 500 pages of other stuff. Then, on page 501, the story mentions "Barnaby" again.
- The Human: Your brain is like a sponge that slowly leaks water. By page 501, you might vaguely remember Barnaby, but you have to work a little to recall who he is. This takes mental energy.
- The AI: The AI has a "perfect replay button." It can instantly recall the exact word "Barnaby" from page 1, no matter how many words are in between. It never forgets.
- The Result: Because the AI never forgets, it predicts the return of "Barnaby" too easily. Humans, who struggle to remember details from 500 pages ago, get surprised or confused. The AI's perfect memory makes it look like it understands the story better than a human, but it's actually cheating by remembering everything perfectly.
Why "Superhuman" is a Bad Thing Here
In most AI applications, being superhuman is great. We want a doctor AI to know more than a human doctor, or a chess AI to be better than a grandmaster.
But in cognitive science (the study of how the mind works), being superhuman is a problem. If you want to build a robot that acts like a human, you don't want a robot that is a genius; you want a robot that makes the same mistakes and has the same limitations as a human. If the AI is too perfect, it stops being a useful model for understanding how our brains actually work.
How Do We Fix It?
The authors suggest we need to build "human-like" AIs by intentionally limiting them:
- Limit the Reading: Instead of training the AI on trillions of words, maybe we should only give it the amount of language a child hears (about 100 million words). This would prevent it from memorizing obscure facts humans don't know.
- Make it Forget: We need to change the AI's architecture so it forgets old words over time, just like our brains do. Maybe we should use different types of computer networks that naturally "decay" or lose information, rather than the current type that remembers everything perfectly.
- Change the Goal: Instead of training the AI to always guess the exact right word, we could train it to guess a range of plausible words. If the sentence is about Elvis, the AI should be okay with guessing "Tupelo," "Memphis," or "Nashville," rather than being 100% sure it's Tupelo.
The Bottom Line
To truly understand how humans read and predict language, we need to stop trying to make AI smarter and start trying to make it more "human." We need to give it a smaller memory, a shorter attention span, and a bit of forgetfulness. Only then will it stop being a super-genius and start acting like a real person.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.