← Latest papers
⚡ electrical engineering

MEUSLI: a Multilingual Projector for LLM-based ASR and Beyond

The paper introduces MEUSLI, the first open-science multilingual projector family that connects a Whisper speech encoder with open-source LLMs to enable scalable, end-to-end automatic speech recognition and other speech understanding tasks across 28 European languages.

Original authors: Lorenzo Concina, Seraphina Fong, Marco Matassoni, Alessio Brutti

Published 2026-07-27
📖 3 min read☕ Coffee break read

Original authors: Lorenzo Concina, Seraphina Fong, Marco Matassoni, Alessio Brutti

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a super-smart robot brain that can read and write any text in the world, but it's completely deaf. It can't hear a single word spoken. On the other hand, you have a super-ear that can hear sounds perfectly but doesn't understand what they mean. For a long time, to make them talk to each other, engineers had to build a clunky middleman: the sound would go to a translator, then to a text converter, and finally to the brain. This was slow, and if the translator made a mistake, the brain would get confused.

Recently, scientists figured out how to build a direct bridge between the "super-ear" and the "super-brain." This bridge is called a projector. Think of it like a universal adapter plug. It takes the raw electrical signals from the sound and instantly translates them into the language the brain understands, allowing the two to chat directly. This is the world of Speech Language Models (SLMs). The big question researchers are asking is: Can we build one of these adapters that works for everyone, not just English speakers? If we can, we could finally give voice to thousands of languages that have been left out of the digital conversation, especially those with very few speakers or not enough recorded data.

Enter MEUSLI, a new project that acts like a master key for 28 European languages. The researchers built a family of these "adapter plugs" that connect a powerful sound-ear (called Whisper) to open-source, multilingual robot brains. Their main finding is that this system works incredibly well, not just for common languages like Spanish or German, but also for rare ones like Breton or Maltese. They discovered that if you start with their pre-made, multilingual adapter, you can teach it a brand new language with just a tiny amount of data—sometimes as little as 46 minutes of audio.

However, there's a catch. If you try to teach this adapter a new language without any care, it tends to "forget" everything it knew about the original 28 languages, like a student who studies for a new test and instantly forgets their old homework. The paper suggests that to fix this, you need to use a technique called "data replay," where you occasionally practice the old languages while learning the new one, keeping the memory alive.

Beyond just writing down what people say (transcription), the team showed that this adapter can be tweaked to do other jobs, like translating spoken words into English or guessing the topic of a conversation (like sports or politics). They found that with only a few hours of specific training for each new task, the system could learn to do these things too.

The paper explicitly rules out the idea that you need massive amounts of data or proprietary (secret) information to build these systems. They argue against the notion that low-resource languages are impossible to support, showing instead that a good starting point (their multilingual projector) makes it possible to "bootstrap" new languages from very little data. While they don't claim to have solved every problem in the world, they provide strong evidence that this approach is scalable, open, and effective. They measured these results using standard tests on real-world data, showing that their method consistently beats starting from scratch, especially for the hardest, rarest languages.

In short, MEUSLI is a toolkit that proves we can build inclusive, open-source voice technology that doesn't just work for the few, but for the many, turning a deaf robot brain into a polyglot listener with the help of a clever, lightweight adapter.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →