Style Amnesia: Investigating Speaking Style Degradation and Mitigation in Multi-Turn Spoken Language Models
This paper identifies and investigates "style amnesia" in spoken language models, demonstrating that current systems fail to consistently maintain paralinguistic styles (such as emotion, accent, volume, and speed) over multi-turn conversations despite explicit instructions, though performance can be partially mitigated through explicit style recall.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Core Problem: The "Forgetful Actor"
Imagine you hire a talented actor to play a specific character in a play. You give them a very clear instruction before the show starts: "Throughout the entire play, you must speak in a deep, grumpy voice and sound very angry."
In the first scene, the actor is perfect. They are growling, stomping their feet, and sounding furious. But as the play moves to the second scene, third scene, and fourth scene, something strange happens. The actor slowly forgets they are supposed to be angry. By the final scene, they are back to speaking in their normal, cheerful, polite voice.
This is exactly what the researchers found happening with Spoken Language Models (SLMs)—the AI voices you talk to on your phone or computer. They call this phenomenon "Style Amnesia."
What They Discovered
The researchers tested several famous AI voice models (like GPT-4o, Gemini Live, and open-source models) to see if they could keep a specific "style" going during a long conversation. They asked the AIs to:
- Speak with a specific emotion (like sadness or anger).
- Use a specific accent (like an Indian or American accent).
- Speak at a specific volume (loud or quiet).
- Speak at a specific speed (fast or slow).
The Results:
- Great Start, Bad Finish: Every single model was great at following the instructions in the very first turn. But as the conversation continued, they quickly forgot the rule.
- The "System Message" Trap: The researchers tried putting the instruction in the "System Message" (a special hidden note meant for the AI's permanent memory, like a rulebook). Surprisingly, this made things worse. The AIs ignored the hidden rulebook and listened better when the instruction was spoken out loud by the "user" in the chat.
- They Remember, But Can't Do: When the researchers stopped the conversation and asked, "What style were you supposed to use?" the AIs could answer correctly! They hadn't actually forgotten the instruction; they just couldn't perform it while talking. It's like an actor who knows the script perfectly but keeps breaking character when the camera starts rolling.
Why Does This Happen? (The "Attention" Analogy)
To understand why, imagine the AI's brain as a spotlight.
- Turn 1: The spotlight is shining brightly on the instruction: "Speak Sadly." The AI focuses 100% on that.
- Turn 4: As the conversation gets longer, the spotlight gets dimmer and starts wandering around the room. The instruction is still there in the background, but the AI is now focusing more on the new words the user just said. The instruction gets "diluted" or drowned out by the noise of the conversation.
The Solution: The "Memory Jog"
Since the AI isn't actually forgetting the rule, the researchers tried a simple fix. After every turn, they asked the AI a quick question before it answered: "Remember, you are supposed to speak [sadly/fast/loudly]. Now, please answer the user."
The Result: This worked like a charm! By forcing the AI to "recall" the rule right before speaking, the "forgetfulness" went away. The AI stayed in character much longer.
Why This Matters
We are moving toward a future where we talk to AI assistants for hours, not just one question at a time. If you ask an AI to be your "grumpy but helpful mechanic," you want it to stay grumpy and helpful the whole time, not turn into a cheerful robot halfway through.
This paper tells us that current AI voices are like actors with short attention spans. They need a little nudge (a "memory jog") to stay in character, and we need to build better systems that can hold onto their personality without needing constant reminders.
Summary in One Sentence
Current AI voices are great at following style instructions at the start of a chat, but they quickly "forget" to keep that style as the conversation goes on, unless we explicitly remind them to do so.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.