← Latest papers
🧠 neuroscience

The cortex encodes speech timing as departure from expectation across multiple timescales

This study demonstrates that the human cortex actively encodes speech timing variability as meaningful information across multiple timescales, specifically representing the deviation of syllable and phoneme durations from expectations rather than treating such variability as noise to be discarded.

Original authors: Molinaro, N., Perez-Navarro, J.

Published 2026-09-03
📖 4 min read☕ Coffee break read

Original authors: Molinaro, N., Perez-Navarro, J.

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

When we listen to a story, our brains do not just hear words; they track the rhythm of the voice. For decades, scientists believed the brain solved the problem of speech timing by locking onto a steady beat, much like a drummer keeping time with a metronome. In this view, the natural variations in how long we hold a vowel or pause between words were seen as background noise, a kind of static that the brain had to filter out to find the underlying regularity. The assumption was that if speech were perfectly regular, our understanding would be even clearer. But this idea leaves a puzzle: real speech is rarely perfectly regular, and when researchers artificially force speech into a robotic, perfectly timed rhythm, people actually understand it worse, especially in noisy environments. This suggests that the irregularity itself might be doing something important, perhaps carrying information that a steady beat cannot.

A team of researchers at the Basque Center on Cognition, Brain and Language set out to test whether the brain treats these timing variations as noise to be ignored or as a signal to be read. They recorded the brain activity of twenty-five native Spanish speakers as they listened to twenty minutes of natural, spontaneous storytelling. Spanish was chosen because it is often described as a language with a steady, syllable-based rhythm, making it the strictest test case for the idea that the brain prefers regularity. The researchers built a computer model to track three different levels of timing expectations: when the next syllable should start, when the next sound within a word should start, and how long a specific sound should last based on what came before. They then compared these expectations against the actual timing of the speech and the electrical activity in the listeners' brains.

The results overturned the old idea that the brain tries to smooth out irregularities. Instead, the brain was found to be encoding the exact moment a sound departed from what was expected. The researchers discovered that the brain does not simply measure how long a sound lasts in milliseconds. Rather, it calculates how surprising the duration is. If a speaker holds a sound for longer than the listener's brain predicted, or shorter, that difference is what triggers a specific response in the auditory cortex. This response was not just a general reaction to sound; it was a precise calculation of the gap between expectation and reality. The study showed that this mechanism works across different speeds of speech, from the slow rhythm of syllables down to the tiny, split-second timing of individual sounds.

Crucially, the brain also pays attention to the direction of the surprise. It distinguishes between a sound that arrives too early and one that arrives too late. These two types of deviations trigger different brain responses at different times, suggesting the brain is not just measuring the size of the error but the nature of it. When the researchers looked at the raw length of sounds versus the surprise of that length, they found that the brain ignored the raw length once it had accounted for the surprise. In other words, the brain does not care that a sound lasted 100 milliseconds; it cares that it lasted 100 milliseconds when the brain expected it to last 40. This finding was confirmed by a behavioral experiment where listeners struggled significantly more to understand sentences when the natural timing was flattened into a perfect, robotic rhythm.

The study also clarified how this timing prediction works alongside the brain's ability to predict what words will come next. The researchers found that the brain handles the timing of events and the meaning of words as separate, parallel processes. The brain predicts when a sound will happen independently of what that sound means. This means that the brain is constantly running a clock that is updated moment by moment, adjusting its expectations based on the immediate past and then recording the deviation when the next sound arrives. This deviation is not discarded as an error; it is retained as a piece of information.

These findings suggest that the irregularity of natural speech is not a flaw the brain must overcome, but a feature it exploits. The brain uses the moment-to-moment departures from expectation to build a rich, detailed understanding of the speech stream. By treating timing as a variable signal rather than a fixed beat, the brain can adapt to the fluid, unpredictable nature of human conversation. The study concludes that the listening brain does not just recover a hidden regularity from a messy signal; it reads the mess itself, finding meaning in the very deviations that were once thought to be noise.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →