← Latest papers
⚡ electrical engineering

Throat and acoustic paired speech dataset for deep learning-based speech enhancement

This paper introduces the Throat and Acoustic Paired Speech (TAPS) dataset, comprising paired recordings from 60 Korean speakers and an optimal alignment method, to address the lack of standard resources for deep learning-based speech enhancement using throat microphones in high-noise environments.

Original authors: Yunsik Kim, Yonghun Song, Yoonyoung Chung

Published 2026-04-23
📖 4 min read☕ Coffee break read

Original authors: Yunsik Kim, Yonghun Song, Yoonyoung Chung

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to have a conversation in the middle of a roaring subway train or a busy factory floor. It's nearly impossible to hear each other because the background noise is so loud.

To solve this, scientists have developed a special kind of microphone that you wear on your throat. Instead of listening to sound waves traveling through the air (like a normal microphone), this "throat mic" feels the vibrations of your vocal cords buzzing against your skin. It's like putting your ear against a wall to hear a party happening in the next room; you can feel the music clearly even if the hallway is noisy.

The Problem: The "Muffled" Signal
However, there's a catch. When sound travels through your skin and muscles to reach that throat mic, it gets muffled. Think of it like looking through a thick, foggy window. You can see the shapes of people, but the fine details (like the sharp edges of letters or the crisp "s" and "f" sounds) get lost. The throat mic captures the "body" of your voice but misses the high-pitched, crisp details that make speech clear and intelligible.

The Solution: A New Recipe Book (The TAPS Dataset)
For a long time, researchers trying to fix this "muffled" voice with Artificial Intelligence (AI) were stuck. They didn't have a good "recipe book" to teach the AI how to turn a muffled throat recording into a clear, normal voice. They were trying to bake a cake without a list of ingredients or a picture of the finished product.

This paper introduces a new, massive recipe book called the TAPS Dataset.

  • What is it? It's a collection of 6,000 pairs of recordings. For every sentence, they recorded it twice at the exact same time: once with the throat mic (the muffled version) and once with a high-quality normal microphone (the clear version).
  • Who made it? 60 native Korean speakers read 100 sentences each in a soundproof room.
  • Why is it special? The researchers didn't just dump the files together. They realized that the throat mic and the normal mic don't always start and stop at the exact same millisecond (like two runners starting a race slightly out of sync). They developed a clever "synchronization dance" to line up the two recordings perfectly before feeding them to the AI.

The Experiment: Teaching the AI
The researchers used this new dataset to train three different types of AI "chefs" to see which one could best restore the lost voice details.

  1. The "Masking" Chef (TSTNN): This chef tried to clean up the existing muffled voice by covering up the bad parts. It was okay, but it couldn't invent the missing sounds.
  2. The "Mapping" Chefs (Demucs & SE-conformer): These chefs were smarter. Instead of just cleaning, they learned to predict and invent the missing high-frequency sounds (like the "s" and "sh" sounds) based on the context of the muffled voice.

The Results
The "Mapping" chefs won. They were able to take the fuzzy throat recording and reconstruct a clear, natural-sounding voice that was almost as good as the original normal microphone recording. The AI learned to fill in the "foggy window" with the missing details.

Why This Matters
This paper is a big deal because it provides the first standard, high-quality "training ground" for this specific technology.

  • For the Future: Now, researchers all over the world can use this same dataset to build better AI. This means we could soon have wearable devices that let you talk clearly in a hurricane, a construction site, or a crowded train, without anyone else hearing you.
  • Beyond Speech: This technology could also help people who have lost their voices or cannot speak out loud to communicate silently by just thinking (or vibrating their throat), turning those vibrations into clear speech.

In short, the authors built the perfect "gym" (the dataset) and the perfect "weights" (the alignment method) to train AI to turn a muffled whisper into a crystal-clear shout, even in the noisiest places on earth.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →