← Latest papers
🧬 biology

A 1000-hour EEG-EMG-audio dataset of Japanese speech production

This paper introduces a publicly available, 1020-hour multimodal dataset comprising time-synchronized scalp EEG, facial EMG, and audio recordings from three native Japanese speakers during overt speech, collected across multiple devices and sessions to support research in speech decoding, artifact modeling, and EEG representation learning.

Original authors: Motoshige Sato, Ilya Horiguchi, Masakazu Inoue, Kenichi Tomeoka, Eri Hatakeyama, Yuya Kita, Atsushi Yamamoto, Ippei Fujisawa, Shuntaro Sasai

Published 2026-06-02
📖 4 min read☕ Coffee break read

Original authors: Motoshige Sato, Ilya Horiguchi, Masakazu Inoue, Kenichi Tomeoka, Eri Hatakeyama, Yuya Kita, Atsushi Yamamoto, Ippei Fujisawa, Shuntaro Sasai

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ⚕️ This is an AI-generated explanation of a preprint that has not been peer-reviewed. It is not medical advice. Do not make health decisions based on this content. Read full disclaimer

Imagine you want to teach a computer to understand how the human brain works when we speak. To do this, you need a massive library of recordings that capture not just the sound of the voice, but the electrical whispers of the brain and the tiny muscle twitches of the face, all happening at the exact same time.

This paper introduces JapanEEG, a massive new "library" containing 1,020 hours of such recordings. It's like building a high-definition, multi-angle movie of the brain in action, specifically focused on Japanese speakers talking out loud.

Here is a breakdown of what they did, using simple analogies:

1. The "Three-Camera" Setup

Usually, scientists might use one type of sensor to record brain activity. This team, however, used three different "cameras" (EEG systems) simultaneously on the same people:

  • The Ultra-High-Def Camera: A system with 128 sensors (g.Pangolin) that acts like a 4K camera, capturing incredibly detailed electrical signals.
  • The Wide-Angle Cameras: Two other systems (g.SCARABEO and eego™sports) with 62–63 sensors each, acting like standard wide-angle lenses to see the whole brain.

By using three different "cameras" on the same subjects over several months, they created a dataset that helps researchers figure out how to translate data from one type of machine to another, solving a common headache in brain research.

2. The "Three-Lens" Recording

To make sure the data is clean and useful, they didn't just record the brain. They recorded three things at once, synchronized perfectly like a multi-track audio recording:

  • The Brain (EEG): The electrical activity of the mind.
  • The Face (EMG): Tiny electrical signals from the lips and eyes. Think of this as a "noise filter" lens. When we speak, our face muscles move, creating electrical static that can muddy the brain signal. By recording the face separately, researchers can later subtract this "noise" to see the pure brain signal.
  • The Voice (Audio): The actual sound of the speech.

3. The "Open-Book" Script

The participants didn't just read a few fixed sentences. They engaged in open-vocabulary speech, which is like reading from a vast, endless library rather than a script with only 10 lines.

  • They read from novels, non-fiction books, comics, and even played text-based video games while reading the dialogue out loud.
  • They also did "covert speech" (imagining speaking) and "listening" tasks.
  • This variety is crucial because it captures the brain doing many different types of "speech work," not just repeating the same phrase.

4. The "Quality Control" Check

Before releasing this data, the team performed a "sound check" to ensure the recordings were real and high-quality.

  • The "Fingerprint" Test: They looked at the electrical "fingerprint" (Power Spectral Density) of the brain waves. They confirmed it looked like a healthy brain (showing the expected drop in power as frequency increases) and that the "static" from power lines was successfully removed.
  • The "Reaction" Test: They checked if the brain reacted to the speech tasks. They found that the brain showed specific, timed electrical spikes (Event-Related Potentials) exactly when the participants started speaking or listening. These reactions looked like organized waves moving across the head, proving the signals came from the brain and not random electrical noise.

5. Why This Matters (According to the Paper)

The authors state that this dataset is a foundation for future work.

  • For Speech Decoding: It helps build better tools to translate brain signals into speech (useful for people who cannot speak).
  • For General Brain Research: Because the data is so large, comes from multiple devices, and includes face muscles and audio, it allows scientists to study how to clean up brain signals, how to train AI models on brain data, and how to make these models work across different people and machines.

In short: The authors have built a massive, high-quality, multi-angle "movie" of the brain speaking Japanese. They have proven the movie is clear and in focus, and they have made the raw film reels available for anyone to study, helping to train the next generation of brain-computer interfaces and signal-processing tools.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →