From Exposure to Expectation: Frequency, Surprisal, and Language Across Development in Spanish
This study demonstrates that while cumulative lexical frequency is a strong predictor of when children acquire Spanish words, contextual surprisal primarily explains moment-to-moment reading processing difficulties in adults, suggesting these factors play distinct roles across different stages of language development.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Language is a river that flows through our lives, carrying us from the first babble of infancy to the complex, rapid-fire conversations of adulthood. For decades, scientists have tried to understand how this river shapes our minds, asking two distinct but related questions. The first is about the beginning: why do children learn some words, like "dog" or "milk," much earlier than others, like "justice" or "whisper"? The second is about the journey: once we know a word, why does it sometimes feel easy to read or hear, and other times feel like a stumble, even when the word itself hasn't changed?
To answer these questions, researchers often look at two different kinds of statistical clues hidden in the language we hear. One clue is frequency, which is simply a count of how often a child hears a specific word in their daily life. The more often a word appears, the more familiar it becomes. The other clue is a measure of surprise. Imagine a sentence where you are waiting for the next word. If the sentence makes you expect a certain word, that word feels predictable. If the sentence leads you to expect something else, and a different word appears, that word carries a high "surprise" value. In the world of modern science, researchers use powerful computer programs, known as large language models, to calculate this surprise. These programs are trained on vast amounts of text and can guess the next word in a sentence with remarkable accuracy, allowing scientists to measure exactly how unexpected a word is in a given context.
For a long time, it was assumed that these two clues—how often we hear a word and how surprising it is—worked together in the same way, whether we were talking about a child learning their first vocabulary or an adult reading a novel. The idea was that the same rules of language learning and processing applied across the board. However, a new study by Francisco Portillo López at the Universidad de Navarra challenges this assumption. By examining the Spanish language through two different lenses, the research suggests that the rules for learning a word are fundamentally different from the rules for processing a word once it is already known.
The researcher began by looking at the earliest stages of language, focusing on 225 common Spanish nouns. He wanted to know if the computer's measure of surprise could help explain why some words are learned early and others late. To do this, he gathered data on how often these words appeared in the speech of caregivers talking to young children, a source known as child-directed speech. He also used three different computer models to calculate the surprise value of each word based on the context in which it appeared. The results were clear and consistent: the frequency of a word was a powerful predictor of when a child would learn it. Words that appeared often in the child's environment were learned much earlier. In contrast, the measure of surprise offered almost no help. Even when the computer calculated how unexpected a word was in a specific sentence, this information did not explain why a child learned that word at a particular time, once the simple fact of how often the word was heard was taken into account.
To ensure this wasn't just a quirk of how the computer was asked to guess the words, the researcher tested the computer in different ways. He tried using standardized sentence frames, where every word was placed in the same simple structure, and he also tried using real, natural conversations between parents and children. In both cases, the pattern held. The computer's sense of surprise, which is designed to measure how well a context predicts a specific word, did not seem to matter for the timing of early word learning. The only thing that reliably predicted when a child would acquire a new word was how many times they had encountered it before. This finding supports a view of learning where repetition and cumulative exposure are the primary drivers, rather than the moment-to-moment puzzle-solving of predicting what comes next.
The story takes a sharp turn when the researcher shifts focus from children to adults. In the second part of the study, he examined how adult Spanish speakers read text. Using data from an eye-tracking study, where cameras recorded exactly where and how long people looked at words on a page, he tested the same computer model's measure of surprise. This time, the results were completely different. For adult readers, the measure of surprise was a strong and reliable predictor of how long they paused on a word. When a word was less predictable given the words that came before it, adults spent more time looking at it, indicating that their brains were working harder to process the unexpected input. This effect remained strong even after accounting for how common the word was and how long it was. In the adult world of skilled reading, the surprise value of a word matters deeply.
The contrast between these two findings reveals a significant split in how language works. For a child building a vocabulary from scratch, the sheer volume of exposure to a word is what cements it in memory. The context in which the word appears matters less than the fact that the child has heard it many times. It is as if the child is collecting stones, and the more times they pick up a specific stone, the more familiar it becomes. For an adult, however, the mind is already full of these stones, organized into a vast network. When an adult reads, they are constantly making predictions about what comes next. If the text breaks that prediction, the brain registers a moment of friction, and the eyes slow down to resolve the mismatch.
This distinction suggests that the computer models used to study language are not failing; rather, they are capturing different aspects of the human experience depending on the stage of life being observed. The same mathematical measure of surprise that is nearly useless for predicting when a child learns a word is essential for understanding how an adult reads it. It implies that the brain does not use a single, static method for handling language. Instead, the importance of statistical clues shifts as we develop. Early on, the brain relies on the accumulation of experience. Later, once a robust system is in place, the brain becomes highly sensitive to the immediate context and the expectations it generates.
The study also offers a cautionary note for how we evaluate artificial intelligence. A computer program might be excellent at mimicking the reading patterns of an adult, making it seem like a perfect model of human language. But this same program might fail to capture the mechanics of how a child actually learns. If we only test these models against adult behavior, we might miss the fact that they do not reflect the learning process of a developing mind. The research suggests that to truly understand language, we must look at both the child and the adult, recognizing that the rules of the game change as the player grows. The journey from exposure to expectation is not a straight line where the same factors apply at every step; it is a path where the weight of frequency and the power of prediction shift, defining different chapters in the story of how we learn to speak and read.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.