The Unheard Alternative: Contrastive Explanations for Speech-to-Text Models
This contribution presents the first method for generating contrastive explanations in speech-to-text models by analyzing how input spectrogram features influence the selection between alternative outputs and demonstrates their effectiveness through a case study on gender assignment in speech translation.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Idea: Asking "Why this and not that?"
Imagine you ask an intelligent assistant to translate a sentence from English into Italian. It says, "I am curious" becomes "Sono curiosa" (using the feminine form).
Normally, if you ask an AI, "Why did you say that?", it might point to the entire sentence and say, "Because of the audio I heard." This is like a detective saying, "The suspect was in the room," without telling you what they saw that made them think it was the suspect.
This paper introduces a new way to ask the question: "Why did you choose curiosa (feminine) instead of curioso (masculine)?"
The authors call this a contrastive explanation. Instead of just explaining the answer, it explains the choice between two options.
The Problem: The "Fuzzy Map"
The researchers investigated how current AI models explain their decisions. They found that existing methods create a "heat map" over the audio recording (a spectrogram).
- The old way: When the AI says "curiosa," the old method highlights the parts of the audio that helped it recognize the word "curious" in general. It is like highlighting the whole kitchen because you made a sandwich. It doesn't tell you which specific ingredient made it a turkey sandwich instead of a ham sandwich.
- The problem: These maps are too broad. They miss the tiny, specific audio clues (like voice pitch) that tell the AI whether the speaker is male or female.
The Solution: A New "Scorecard"
To fix this, the team developed a new method based on an existing tool called SPES. Imagine SPES as a way to play "hide-and-seek" with the audio. It covers up small pieces of sound and sees if the AI changes its mind.
However, to make this work for "Why A instead of B?", they had to invent two new rules:
- The "Whole Word" Rule: AI models usually speak in tiny sound fragments (subwords), like building blocks. The old methods simply added up the blocks. The new method is smarter; it checks whether these blocks actually form a complete word or just the beginning of a longer word. It is like making sure you don't count the word "pre-" as a complete word just because it is part of "prepare."
- The "Relative Score" (the secret spice): This is the most important part.
- The old score (Difference): It asks, "How much did covering up this sound reduce the chance of saying curiosa?" If the AI was already 99% sure it was curiosa, covering up the sound hardly changes anything. The score stays low, and the map looks boring.
- The new score (Relative): It asks, "How much did covering up this sound help curioso (the other option) catch up?" It compares the two options directly.
- The Analogy: Imagine a race between a Ferrari (the AI's choice) and a bicycle (the alternative).
- The old score looks at how much the Ferrari slows down when you put a rock in its tire. It hardly slows down, so the rock doesn't seem important.
- The new score looks at how much the bicycle speeds up when you remove a rock from its path. If the bicycle suddenly catches up, that rock was the decisive factor.
The Experiment: Gender in Translation
To test this, the researchers used a specific scenario: Gender assignment.
In languages like Italian, French, and Spanish, adjectives change depending on gender (e.g., curioso vs. curiosa). English does not do this, so the AI must guess the speaker's gender based on their voice (pitch, timbre, etc.).
They tested their new method on audio clips where the AI had to choose between male and female translations.
The Results:
- The old method: Generated maps that looked almost identical to the "broad" maps. It could not distinguish between "why this word" and "why this gender."
- The new method: Generated maps that highlighted very specific parts of the audio.
- When they blocked the specific audio properties that the new method had identified, the AI stopped saying "female" and switched to "male" in over 70% of cases.
- This proved that the method successfully found the exact "voice clues" the AI used to guess the gender.
The Catch (Limitations and Ethics)
The paper is very honest about the drawbacks:
- The "Default" Bias: The AI seems to assume everyone is male unless it hears strong evidence that they are female. When the researchers tried to force the AI from "male" to "female" by blocking audio, it didn't work as well. This suggests the AI has a built-in bias toward the male default, likely because it was trained on more male voices than female voices.
- The Binary Problem: The study only looked at "Male" vs. "Female." The authors note that this ignores non-binary speakers. If a person's voice does not fit the typical "male" or "female" pattern, the AI might get confused or provide a translation that does not fit them.
- Ethical Warning: Using voice features to estimate gender can be difficult. Transgender people or individuals with speech disorders might receive "wrong" translations because their voices do not match the AI's old-fashioned expectations.
Summary
This paper didn't just build a better map; it built a magnifying glass.
- Before: We could see where the AI looked to make a decision, but it was a blurry, general view.
- Now: We can see exactly which specific sound made the AI choose "she" instead of "he."
This helps researchers understand how AI models "listen," and helps them recognize when the AI makes unfair assumptions based on gender.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.