← Latest papers
💻 computer science

Omnilingual ASR: Open-Source Multilingual Speech Recognition for 1600+ Languages

This paper introduces Omnilingual ASR, a scalable, open-source family of speech recognition models that leverages a 7B-parameter self-supervised architecture and a massive, diverse training corpus to support over 1,660 languages, including more than 500 previously unserved, while enabling communities to easily adapt the system to new languages with minimal data.

Original authors: Marta Costa-jussa, Omnilingual Team, Gil Keren, Artyom kozhevnikov, Yen Meng, Christophe Ropers, Matthew Setzler, Skyler Wang, Ife Adebara, Michael Auli, Can Balioglu, Kevin Chan, Chierh Cheng, Joe Ch
Published 2026-07-30
📖 4 min read☕ Coffee break read

Original authors: Marta Costa-jussa, Omnilingual Team, Gil Keren, Artyom kozhevnikov, Yen Meng, Christophe Ropers, Matthew Setzler, Skyler Wang, Ife Adebara, Michael Auli, Can Balioglu, Kevin Chan, Chierh Cheng, Joe Chuang, Caley Droof, Mark Duppenthaler, Paul-Ambroise Duquenne, Alexander Erben, Cynthia Gao, Gabriel Mejia Gonzalez, Kehan Lyu, Sagar Miglan, Vineel Pratap, Kaushik Ram Sadagopan, Safiyyah Saleem, Arina Turkatenko, Albert Ventayol-Boada, Zheng Xin Yong, Yu-An Chung, Jean Maillard, Rashel Moritz, Alexandre Mourachko, Mary Williamson, Shireen Yates

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a magical translator that can instantly understand and write down what anyone is saying, no matter what language they speak. For a long time, this magic only worked for the most popular languages, like English, Spanish, or Mandarin. If you spoke a smaller, less common language, the translator would just stare blankly or guess wildly. This is the world of Automatic Speech Recognition (ASR): the technology that turns spoken words into text. Think of it like a super-fast scribe that listens to a conversation and types it out. The big challenge has always been that there are over 7,000 languages in the world, but most of these scribes only know a tiny fraction of them. Learning a new language usually requires a massive library of recorded conversations and a team of experts to teach the computer, which is expensive and slow. But what if we could build a scribe that learns the patterns of how all languages work, so it can pick up a new one almost instantly, just by hearing a few sentences? That is the big question this paper tackles: how do we make speech technology fair and accessible for everyone, not just the few who speak the "big" languages?

Enter Omnilingual ASR, a new open-source project from Meta that is like a super-powered, polyglot scribe designed to speak over 1,600 languages. The team didn't just try to teach the computer more words; they built a giant, diverse "brain" that learned from a massive collection of speech data, including recordings from communities that have never been heard by AI before. They created a system that comes in different sizes: a huge, super-accurate version (7 billion parameters) for powerful computers, and a tiny, lightweight version (300 million parameters) that can run on a simple phone. The most exciting part is how they handle languages the computer has never seen before. Instead of needing a whole library of data to learn a new language, this system can learn from just 10 short examples of someone speaking and writing. It's like showing a master chef a picture of a new fruit and a recipe for a dish made with it; suddenly, they can cook that dish without needing to taste the fruit a thousand times first.

The researchers found that this new system is a game-changer, especially for languages with very little data available. When they tested it, the model achieved a "Character Error Rate" (a score measuring how many letters it got wrong) of under 10% for more than 1,200 languages. In fact, for over 500 languages, this is the very first time any speech system has ever been able to transcribe them. The paper shows that their new model beats previous champions, like Whisper and USM, especially on those rare, long-tail languages. While older systems would get confused and fail on these languages, Omnilingual ASR kept its cool. They also discovered that if you want to make the system even better for a specific, tiny language, you can fine-tune the smaller versions with very little computing power and just a few hours of data, getting results that rival the massive models.

However, the paper is careful to note that this isn't a magic wand that solves everything perfectly. The system is still learning, and while it can handle unseen languages with just a few examples (a "zero-shot" capability), it doesn't perform quite as well as a system trained specifically on that language with tons of data. The authors suggest that for the most critical uses, like medical or legal transcription, humans should still double-check the work. They also emphasize that this technology is meant to be a tool for communities to preserve their languages and build their own archives, not to replace human speakers or force a one-size-fits-all solution. By releasing all their code and data for free, they hope to invite researchers and communities everywhere to join the party, making it easier to build a future where every language has a voice in the digital world.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →