Lost but not erased: Finding traces of a forgotten language in neural speech models
This study demonstrates that traces of a forgotten birth language persist in neural speech models trained on a second language, suggesting that critical-period effects in language acquisition stem from the entrenchment of foundational representations rather than a maturational loss of plasticity.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Language is often thought of as a skill that must be learned within a specific window of childhood, a biological deadline after which the brain loses its ability to absorb new sounds and patterns with the same ease. This idea, known as the critical period, suggests that if a child stops speaking their native tongue early in life, that language is gone forever, erased by the maturation of the brain. Yet, there are people who challenge this simple erasure. International adoptees, children who are taken from their birth families and raised in a new country, often lose the ability to speak or understand their first language completely. As adults, they have no conscious memory of it. However, careful testing has shown that these individuals can relearn the sounds of their birth language much faster than people who never heard it at all. This phenomenon, where forgotten knowledge returns with surprising speed, suggests that the language was not truly deleted, but merely hidden. The question that has long puzzled scientists is why this happens: is it because the human brain has a special, time-limited biological mechanism that preserves these traces, or is it simply a result of how learning systems, whether human or artificial, build their foundations?
To investigate this, researchers turned to artificial intelligence, specifically computer models designed to recognize speech. These models are not biological; they do not grow, age, or mature in the way humans do. Instead, they learn purely through exposure to data, making them a perfect tool to isolate the effects of experience from the effects of biology. The scientists trained these models to understand one language, such as German, and then abruptly switched the training data to a completely different language, like French. This setup mimicked the experience of an international adoptee, forcing the model to abandon its first language and master a second one. As expected, the models quickly lost their ability to recognize the first language, dropping to a level of performance indistinguishable from random guessing, while simultaneously becoming highly proficient in the new language. On the surface, the first language appeared to be gone, just as it is for human adoptees.
However, when the researchers looked inside the model's internal structure, they found that the first language had not vanished. The model's architecture is built in layers, similar to a multi-story building where the lowest floors handle basic details and the upper floors handle complex meanings. The researchers discovered that traces of the lost language persisted, but they were concentrated almost entirely in the lowest, most foundational layers of the network. These early layers are responsible for processing basic sound patterns, the raw building blocks of speech, before they are combined into words or sentences. Even after the model had spent thousands of hours learning the new language, these bottom layers still retained the statistical fingerprints of the original language. The higher layers, which deal with complex grammar and vocabulary, had been completely overwritten, but the foundation remained stubbornly unchanged.
To test if these hidden traces were actually useful, the researchers asked the models to relearn their lost language. The models that had experienced the early exposure were able to relearn the first language significantly faster than models that had never heard it at all. In fact, they reached a high level of accuracy in about fourteen percent fewer steps. This speed advantage proved that the traces were not just passive leftovers; they were functional. The models were essentially using the old, foundational sound maps to build the new language skills, giving them a head start. Crucially, when the researchers swapped out the lowest layers of the fast-learning model with layers from a model that had never learned the first language, the speed advantage disappeared. This confirmed that the benefit came specifically from those early, entrenched layers.
The study also explored how long the initial exposure needed to be to create these lasting traces. They found that the effect was not a simple matter of "the longer, the better." Instead, the traces grew stronger as the model learned the first language for a while, peaked after a certain amount of time, and then began to slowly decline if the training continued too long. This suggests there is an optimal window for these foundational representations to become entrenched, after which they become stable and resistant to change. The researchers argue that this behavior does not require a special biological clock or a maturational deadline. Instead, it appears to be a natural consequence of how learning systems work: once a system builds a stable foundation, it becomes very difficult to completely rewrite that foundation without destabilizing everything built on top of it.
These findings offer a new perspective on the critical period in language learning. Rather than being a biological window that closes, the difficulty of unlearning early experiences may simply be a result of the system's own structure. The early learning creates a durable, low-level scaffold that supports all future learning. When the input changes, the system adapts its higher-level functions to the new language, but the foundational layers remain, preserving a trace of the past. This means that the "savings" effect seen in international adoptees might not be a mystery of human biology, but a general property of learning itself. The language is not erased; it is buried deep in the foundation, waiting to be uncovered when the system is asked to build on it again.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.