← Latest papers
🤖 machine learning

LibriBrain100: One Hundred Hours of Broad and Deep MEG Data for Neural Speech Decoding at Scale

The paper introduces LibriBrain100, a large-scale MEG dataset featuring over 100 hours of neural recordings from naturalistic speech that combines deep within-subject data with broad multi-subject sampling to establish a standardized benchmark and open competition for advancing non-invasive brain-to-text decoding.

Original authors: Francesco Mantegna, Dulhan Jayalath, Gereon Elvers, Tasha Kim, Benjamin Ballyk, Alex Fung, SungJun Cho, Teyun Kwon, Luisa Kurth, Miran Özdogan, Gilad Landau, Pratik Somaiya, Natalie Voets, Mark Woolri
Published 2026-08-27
📖 6 min read🧠 Deep dive

Original authors: Francesco Mantegna, Dulhan Jayalath, Gereon Elvers, Tasha Kim, Benjamin Ballyk, Alex Fung, SungJun Cho, Teyun Kwon, Luisa Kurth, Miran Özdogan, Gilad Landau, Pratik Somaiya, Natalie Voets, Mark Woolrich, Oiwi Parker Jones

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a world where a person who has lost the ability to speak due to paralysis could communicate again, not by typing or blinking, but simply by thinking their words. This is the promise of a brain-computer interface, a technology that reads the electrical and magnetic signals of the brain to decode what a person is saying or trying to say. For years, the most successful versions of this technology have required surgeons to implant electrodes directly onto the brain, a risky procedure that limits its use to a tiny fraction of patients. The scientific community has long sought a non-invasive alternative, one that could read the brain from the outside, perhaps through a helmet or a cap, making the technology safe and accessible for anyone. However, progress in this field has been slow and difficult to measure. Without a shared set of data and a standard way to test new ideas, researchers have struggled to know if their methods are truly improving or just working by chance. To solve this, scientists need a massive library of brain recordings, collected under strict conditions, that everyone can use to build and test their decoding tools.

A team of researchers has now released such a library, a dataset called LibriBrain100, which represents a significant leap forward in the quest to read speech from the brain without surgery. The project focuses on magnetoencephalography, or MEG, a technique that uses incredibly sensitive sensors to detect the faint magnetic fields produced by the brain's activity. While the human brain is a complex organ, the magnetic signals it generates when processing speech are distinct enough to be captured, provided the equipment is sensitive enough and the data is abundant. The challenge has always been that collecting this data is difficult and time-consuming. To get a clear picture of how the brain handles language, researchers need hours of recording from the same person listening to speech, as well as recordings from many different people to see how the brain's language centers vary from one individual to another. Before this new release, the largest available collection of such data was limited, often focusing on just one person or a few hours of recording, which made it hard to train the powerful computer models needed for real-world applications.

The new LibriBrain100 dataset changes the scale of the problem entirely. It contains more than one hundred hours of high-quality brain recordings, a volume more than double that of the previous largest collection. The core of this dataset comes from a single volunteer who listened to natural, continuous speech for approximately eighty hours. This person listened to the complete collection of Sherlock Holmes stories, a set of phonetic exercises designed to cover a wide range of speech sounds, and a series of podcast narratives covering diverse topics. This deep focus on one individual allows researchers to study the brain's language processing with a level of detail that was previously impossible, capturing the subtle ways the brain responds to the same speaker over a long period. To address the need for understanding how these findings apply to other people, the researchers also recorded about forty minutes of data from thirty-two additional volunteers, all listening to a portion of the Sherlock Holmes stories. This combination creates a unique resource that is both deep, with massive amounts of data from one person, and broad, with data from many people, allowing scientists to test if a model trained on one person can be adapted to work for another with very little new data.

The researchers used this dataset to test a specific task: word classification. In this experiment, a computer model listens to the brain signals recorded while a person hears a word and tries to guess which word from a list of fifty common words was just spoken. This is a stepping stone toward the ultimate goal of full brain-to-text decoding. When they trained a computer model using the deep data from the single volunteer, the model achieved the best performance ever recorded for this type of non-invasive task. This result validates the quality of the recordings and proves that having a large amount of data from a single person is incredibly valuable for teaching a computer how to read that specific brain. However, the real test for a practical device is whether it can work for a new user without requiring eighty hours of training. The researchers found that by taking a model trained on the deep data and giving it just a small amount of data from a new person—about forty minutes—the model could be fine-tuned to work well for that new person. This suggests that a system could be built where a large, pre-trained model is adapted to a new patient with only a short session of listening, making the technology much more feasible for real-world use.

The dataset also includes recordings from different types of speech to ensure the models are robust. The Sherlock Holmes stories provide a consistent narrative, while the phonetic exercises offer a controlled mix of sounds, and the podcasts introduce natural, spontaneous speech with varied topics and speakers. By including these different sources, the researchers can study how the brain processes the sound of speech versus the meaning of speech. They found that the brain signals contain information about both the acoustic properties of the voice and the semantic content of the story. The researchers made all of this data available to the public, along with a software library that makes it easy for other scientists to download and use the recordings. They also launched an open competition where researchers from around the world can test their own decoding methods against the same standard benchmarks. This approach mirrors the success of large datasets in other fields of artificial intelligence, where shared resources have accelerated progress by allowing everyone to build on the same foundation.

While the results are promising, the researchers are careful to note the current limitations. The data was collected while people were passively listening to stories, which is different from the active effort of trying to speak or imagine speaking, a state that is crucial for a communication device for someone who is paralyzed. The team is actively collecting data on inner speech and plans to release it in the future. Furthermore, the current success is limited to identifying words from a small list of fifty, rather than generating full sentences. The researchers did not include a baseline for full brain-to-text decoding in this release because the field is still evolving too rapidly for a single standard to be stable. However, the dataset provides the necessary infrastructure to move toward that goal. By offering a massive, high-quality, and standardized resource, LibriBrain100 removes a major barrier to entry for researchers. It allows the community to focus on developing better algorithms rather than struggling to collect data, with the ultimate hope of creating a non-invasive tool that can restore the ability to communicate to people living with severe paralysis.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →