Neural tracking of surprisal and semantic distance in naturalistic movie viewing
Using fMRI data from participants watching a full-length film, this study demonstrates that the human brain tracks both surprisal and semantic distance as unique predictors of neural activity in bilateral temporal and frontal regions during naturalistic audiovisual language comprehension.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer
When we listen to a story, our brains are not just passive recorders of sound; they are active predictors, constantly guessing what comes next based on what we have already heard. This process relies on two distinct but related mental tools. The first is a sense of probability: how likely is the next word to appear given the words that came before it? If a sentence ends with "The sky is," our brain strongly expects the word "blue," and the arrival of that word feels unsurprising. If the sentence ends with "The sky is," and the next word is "banana," the brain is caught off guard. This feeling of being caught off guard is known as surprisal. The second tool is a measure of meaning, or how closely related the current word is to the ideas we have just processed. Even if a word is not the most statistically likely one, it might still fit perfectly with the theme of the conversation. Understanding how the brain uses these two different types of information—statistical likelihood and semantic relatedness—to make sense of language is a central question in neuroscience, especially when we are immersed in a rich, real-world environment like watching a movie.
A team of researchers set out to explore how the human brain tracks these two specific signals while people watch a full-length film. They recruited twenty adults who had never seen the 2009 movie 500 Days of Summer and asked them to watch the entire film, which lasts just over ninety minutes, while lying inside a functional magnetic resonance imaging scanner. This machine allowed the scientists to watch the brain's activity in real time as the story unfolded. To understand what the brain was doing, the researchers first broke down the movie's audio and visual tracks into thousands of tiny data points. They measured everything from the brightness of the screen and the movement of the camera to the loudness of the dialogue and the frequency of specific words. They then calculated two complex numbers for every word spoken in the film: one representing how surprising that word was based on the previous context, and another representing how semantically distant or close that word was to the words that came before it.
Using a sophisticated computer modeling technique, the researchers built a series of step-by-step predictions to see which of these data points could best explain the patterns of activity they saw in the brain. They started with the simplest factors, such as the visual changes on the screen and the raw volume of the sound. These basic features successfully predicted activity in the back of the brain, where visual processing happens, and in the sides of the brain, where sound is first received. As they added more complex language features, such as how often a word is used in general or how many similar-sounding words exist, the predictions improved, spreading to areas involved in language and memory. Finally, they added the two main variables of interest: the statistical surprise of the words and their semantic distance. The results showed that including these two factors significantly improved the ability to predict brain activity, particularly in the superior temporal gyrus, a region on the sides of the brain known for processing speech, and in the left frontal lobe.
The most critical discovery was that these two factors were not just doing the same job. Even after the researchers accounted for the influence of one, the other still explained unique changes in brain activity. This suggests that the brain is tracking both the statistical probability of what comes next and the conceptual relationship between ideas simultaneously, using two complementary systems. While both signals activated similar regions in the temporal lobes, the statistical surprise of words also showed a unique connection to a specific area in the cerebellum, a part of the brain at the base of the skull often associated with coordination and timing. This finding indicates that the brain's prediction mechanisms are nuanced, handling the "what is likely to happen" and "what fits the meaning" as separate but parallel streams of information.
These results are significant because they move beyond the controlled, quiet environments of traditional laboratory studies. Previous research on language prediction often relied on people listening to isolated sentences or reading text in silence. By using a full movie, the researchers demonstrated that the brain continues to track these subtle linguistic signals even when it is bombarded with complex visual scenes, music, and emotional storytelling. The study confirms that the brain's ability to anticipate language is robust and deeply integrated into how we experience the world, operating effectively even when the input is rich, multimodal, and unpredictable. The work does not claim to have solved the entire mystery of language processing, but it provides strong evidence that our brains are constantly running two different kinds of checks on the words we hear, ensuring we understand not just what is being said, but why it matters in the flow of the story.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.