← Latest papers
⚡ electrical engineering

FiLM-Based Speaker Conditioning of a SpeechLLM for Pathological Speech Recognition

This paper proposes a parameter-efficient approach using Feature-wise Linear Modulation (FiLM) to condition a frozen SpeechLLM on pathological speakers via x-vectors, demonstrating that this method achieves competitive automatic speech recognition performance on Spanish and English pathological data while preserving the model's general conversational capabilities.

Original authors: Fernando López, Santosh Kesiraju, Jordi Luque

Published 2026-06-05
📖 4 min read☕ Coffee break read

Original authors: Fernando López, Santosh Kesiraju, Jordi Luque

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a super-smart robot librarian who is incredibly good at reading standard, clear books. This robot is so talented that it can transcribe normal speech with almost no mistakes. However, when someone with a neurological condition (like Parkinson's or ALS) tries to speak, their voice might be shaky, quiet, or slurred. The robot gets confused, just like a person trying to read a book written in messy, shaky handwriting.

The researchers in this paper wanted to help this robot understand these specific voices without having to retrain the robot from scratch or risk making it forget how to read the clear books.

Here is the simple breakdown of what they did:

1. The Problem: One Size Doesn't Fit All

Standard speech recognition systems are trained on "healthy" voices. When a person with a speech disorder speaks, the robot struggles because the sound patterns are different. Usually, to fix this, engineers would "fine-tune" the robot's brain. But this is risky: if you tweak the brain too much to understand one specific type of messy voice, the robot might forget how to understand clear voices. It's like teaching a chef to make a specific spicy dish so well that they forget how to make a simple salad.

2. The Solution: The "Smart Glasses" (FiLM)

Instead of rewriting the robot's entire brain, the researchers gave it a pair of smart glasses.

  • The Robot (The Base Model): They kept the main robot frozen and untouched. It stays exactly as smart as it was before.
  • The Glasses (FiLM): They added a small, new layer called "FiLM" (Feature-wise Linear Modulation). Think of this as a set of glasses that the robot puts on only when it hears a specific person.
  • How it works: Before the robot listens to a sentence, a small detector figures out "Who is speaking?" (using a digital voice print called an x-vector). If the speaker has a speech disorder, the glasses adjust the robot's internal hearing to match that specific person's voice. If the speaker is healthy, the glasses turn off (become clear), and the robot listens normally.

This way, the robot can adapt to a specific person's unique voice without changing its core knowledge.

3. The Test: Can it still do other things?

The researchers didn't just test if the robot could transcribe words. They also asked it questions about the speaker, like "Is the speaker male or female?" or "How old are they?"

  • The Goal: They wanted to make sure that by giving the robot these "glasses" to understand messy speech, they didn't break its ability to answer general questions.
  • The Result: The robot with the "glasses" got better at understanding the messy speech. Crucially, it also got better at guessing the speaker's gender, even though it wasn't explicitly trained to do that. It seems that by focusing on the speaker's unique voice, the robot became more sensitive to those details.

4. The Catch: It needs a little help

While the "glasses" method worked well, the researchers found that the robot's initial guesses were sometimes a bit "noisy" (like hearing a word but not being 100% sure of the spelling). They had to use a simple "spell-checker" (post-processing) to clean up the final text.

  • Comparison: Other methods that tried to retrain the whole robot's brain made fewer mistakes initially, but they risked making the robot forget how to handle normal voices. The "glasses" method kept the robot safe and flexible, even if it needed a little extra cleaning at the end.

The Bottom Line

The paper shows that you can help a super-smart speech robot understand difficult, pathological voices by giving it a lightweight, adjustable "adapter" (the FiLM glasses) rather than rebuilding its brain.

  • It works: It improves understanding of disordered speech.
  • It's safe: It doesn't ruin the robot's ability to understand normal speech.
  • It's efficient: It only changes a tiny fraction of the robot's brain (about 1.6%), making it a very efficient way to adapt technology for people with speech disorders.

In short, they found a way to give the robot a "personal translator" for specific voices without breaking its original brain.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →