← Latest papers
💬 NLP

SpeakerSleuth: Can Large Audio-Language Models Judge Speaker Consistency across Multi-turn Dialogues?

The paper introduces SpeakerSleuth, a benchmark revealing that while Large Audio-Language Models possess inherent acoustic discrimination skills, they struggle to reliably judge speaker consistency in multi-turn dialogues due to a significant bias toward prioritizing textual coherence over acoustic cues.

Original authors: Jonggeun Lee, Junseong Pyo, Gyuhyeon Seo, Yohan Jo

Published 2026-04-21
📖 4 min read☕ Coffee break read

Original authors: Jonggeun Lee, Junseong Pyo, Gyuhyeon Seo, Yohan Jo

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are watching a movie where the main character, let's call him Bob, is having a long conversation with his friend. In a perfect world, Bob's voice should sound exactly the same in every single sentence he speaks. But what if, halfway through the movie, the actor recording Bob's lines gets tired, or a computer glitch happens, and suddenly Bob sounds like a completely different person? Or maybe he sounds like a slightly different version of himself?

This is the problem SpeakerSleuth is trying to solve.

The Big Idea: The "Voice Detective"

The researchers built a benchmark called SpeakerSleuth (like a detective agency for voices) to test a new kind of AI called a Large Audio-Language Model (LALM).

Think of these LALMs as super-smart robots that can hear audio and read text at the same time. They are being trained to be "judges" for AI voice generators. The question the paper asks is: "Can these robots reliably tell if a character's voice stays consistent throughout a whole conversation, or do they get confused?"

The Three Tests (The Detective's Toolkit)

To test the robots, the researchers created three specific challenges, like a training course for a detective:

  1. The "Spot the Imposter" Test (Detection):

    • The Scenario: You hear a conversation. Is the main character's voice the same person from start to finish, or did someone switch actors?
    • The Result: The robots were terrible at this. Some were so paranoid they thought every conversation had a switch (false alarms). Others were so lazy they thought nothing was wrong, even when the voice clearly changed. They couldn't find a "middle ground."
  2. The "Pinpoint the Mistake" Test (Localization):

    • The Scenario: You know the voice changed at some point. Can you tell exactly which sentence the switch happened?
    • The Result: Even worse. The robots were like a detective who says, "I know a crime happened, but I have no idea who did it or when." They often guessed randomly or flagged the wrong sentences.
  3. The "Best Match" Test (Discrimination):

    • The Scenario: You have three different audio clips of the character. Which one sounds the most like the real Bob?
    • The Result: This is where the robots shined! When they just had to compare and rank voices (like picking the best apple from a basket), they were very good at it. They could hear the subtle differences.

The Big Surprise: The "Text Trap"

Here is the most interesting part of the paper. The researchers gave the robots a hint: they showed them the text of the conversation (what the other people were saying) to help them understand the context.

You would think this would help, right? Like giving a detective a script to follow?
Nope. It made them worse.

  • The Analogy: Imagine a detective who is so obsessed with the story of the crime that they ignore the physical evidence.
  • What happened: When the robots saw the text, they stopped listening to the voices. They thought, "Oh, the story makes sense, so the voices must be right!" Even if the voice suddenly switched from a man to a woman, if the text flowed smoothly, the robot said, "Everything is fine!"

They prioritized the words over the sound. This is called a "modality imbalance"—they care too much about reading and not enough about listening.

Why Does This Matter?

We are entering an era where AI can generate entire movies, podcasts, and conversations with multiple characters.

  • If you are making an AI movie, you don't want the hero to sound like a different person in every scene.
  • If you are building a voice assistant, you don't want it to sound like a robot one second and a human the next.

The paper concludes that while our current AI "judges" are great at comparing two sounds, they are not yet reliable enough to police a whole conversation. They need to be taught to listen more carefully and not get distracted by the text.

Summary in a Nutshell

  • The Goal: Can AI judges spot when a speaker's voice changes in a long conversation?
  • The Findings:
    • They are good at comparing sounds (Discrimination).
    • They are bad at spotting if a whole conversation is consistent (Detection).
    • They are terrible at finding where the mistake happened (Localization).
    • The Trap: If you give them the text script, they ignore the audio and get fooled by the story.
  • The Future: We need to build better "ears" for these AI models so they don't just read the script but actually listen to the voice.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →