← Latest papers
⚡ electrical engineering

BanglaVowelDataset: True AC and Synthetic BC Bangla Vowel Datasets in Speech Information System

This paper introduces the BanglaVowelDataset, comprising 120 recorded air-conducted (AC) and 120 synthetic bone-conducted (BC) Bangla vowels from ten speakers, to address the lack of such resources and facilitate research in noise-robust speech processing, recognition, and speaker identification.

Original authors: Ohidujjaman -, Bejoy Munshi, Mahmudul Hasan, Mohammad Nurul Huda, Suman Ahmmed, Hasan Sarwar, Tetsuya Shimamura

Published 2026-09-07
📖 4 min read☕ Coffee break read

Original authors: Ohidujjaman -, Bejoy Munshi, Mahmudul Hasan, Mohammad Nurul Huda, Suman Ahmmed, Hasan Sarwar, Tetsuya Shimamura

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the world of speech technology, computers are constantly learning to listen. They are trained on vast libraries of recorded voices to recognize words, identify speakers, or understand commands. Most of these libraries, however, are built from sounds captured in the air. When we speak, sound waves travel through the air, hit a microphone, and are converted into digital data. This method, known as air-conduction, is the standard for almost all voice assistants and transcription services. But there is another way sound reaches our ears and recording devices: through the bones of our skull. When we speak, our vocal cords vibrate our skull bones, and these vibrations travel directly to the inner ear and can be picked up by sensors placed on the skin. This is bone-conduction. While air-conducted sound is easily drowned out by the hum of traffic or the chatter of a crowd, bone-conducted sound remains remarkably clear because it bypasses the noisy air. For decades, researchers have struggled to build computers that can understand this bone-based speech, largely because the necessary data simply did not exist for many languages.

A team of researchers from universities in Bangladesh and Japan has now filled a significant gap in this field by creating the first dedicated collection of vowel sounds in the Bangla language, captured through both air and bone. The Bangla alphabet contains twelve distinct vowel sounds, a complexity that makes the language rich but difficult for machines to parse, especially when background noise interferes. Until this work, no dataset existed that offered both the traditional air-recorded versions of these sounds and their bone-conducted counterparts for comparison. The researchers set out to record these sounds from ten native speakers—five men and five women—in a quiet, soundproof room. They used high-quality microphones to capture the air-conducted vowels, ensuring each recording was long enough to be analyzed, ranging from 600 to 780 milliseconds. To ensure the data was clean, they carefully trimmed the silent beginnings and ends of each recording, leaving only the pure, voiced portion of the sound.

The challenge then became how to obtain the bone-conducted data without requiring every participant to wear specialized, rare bone-conduction microphones, which are difficult to find and use in standard research settings. Instead of recording the bone vibrations directly, the team used a sophisticated computer model to simulate them. They took the clean, air-recorded vowels and passed them through a digital filter that mimics how the human skull alters sound. This process effectively stripped away the high-frequency details that air carries but that bones filter out, creating a synthetic version of what the sound would have looked like if it had been recorded through the skull. The result was a paired dataset: 120 real air-conducted vowels and 120 synthetic bone-conducted vowels, perfectly matched to the same speakers and the same twelve vowel sounds.

The researchers then organized this massive collection into a structured library, separating the data into different stages of processing to help other scientists test their own theories. They provided the raw recordings, the trimmed versions, and the computer-simulated bone sounds, all labeled with precise details about the speaker's gender and the specific vowel being spoken. To prove that their simulation was accurate, they compared the mathematical properties of the real air sounds against the synthetic bone sounds. They found that the synthetic bone sounds behaved exactly as expected: they showed a wider range of frequencies and a specific mathematical signature that distinguishes them from air sounds. This confirmed that their computer model successfully recreated the unique acoustic profile of bone-conducted speech.

This dataset is now available for the global research community to use. It offers a unique opportunity to test how well speech recognition systems can handle noise. Because air-conducted speech is easily corrupted by background noise while bone-conducted speech is not, researchers can use this paired data to build systems that switch between the two or combine them to understand speech in chaotic environments like busy streets or crowded rooms. The team has also made the computer code they used to generate these sounds public, allowing others to verify their methods or apply the same techniques to other languages. By providing the first comprehensive look at Bangla vowels through both air and bone, this work opens the door for more robust, noise-resistant speech technology that can function reliably anywhere, not just in a quiet studio.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →