← Latest papers
⚡ electrical engineering

DG^VoiC: Speaker Clustering for Fraud Investigation under Real Call-Centre Conditions

This paper introduces DG^VoiC, a voice clustering framework that leverages anonymized real call-centre audio to identify repeated speakers across customer interactions, achieving high clustering accuracy (up to 100% homogeneity) to enhance insurance fraud investigation by verifying speaker consistency.

Original authors: Muhammad Shakeel Akram, Amal Htait, Abdul Hamid Sadka, Emma Meisingseth, Karishma Jaitly

Published 2026-06-29
📖 4 min read☕ Coffee break read

Original authors: Muhammad Shakeel Akram, Amal Htait, Abdul Hamid Sadka, Emma Meisingseth, Karishma Jaitly

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to solve a mystery in a giant, busy call center. Thousands of people call in every day to report insurance claims. Usually, the police (or in this case, the fraud investigators) look at the words people say or the paperwork they fill out to find liars. But this paper introduces a new tool called DGVoiC that listens to the voices themselves to catch a different kind of trickster.

Here is the simple breakdown of what the paper does, using everyday analogies:

The Problem: The "Masked" Caller

Imagine a fraudster who wants to steal money. They might call in using one name today, and then call in again next week using a completely different name. To a human agent, these look like two different people. But to a voice detective, it's the same person wearing a different mask.

The problem is that call centers are noisy (like a busy coffee shop), and people have different accents or moods. Also, the recordings are "anonymized" (blurred) to protect privacy, so the system can't hear names or addresses, only the sound of the voice.

The Solution: DGVoiC (The Voice Fingerprint Scanner)

The authors built a system called DGVoiC. Think of it as a high-tech "voice fingerprint" scanner that works even in a noisy room.

  1. Cleaning the Audio: First, the system acts like a sound engineer. It mutes out any part of the call where someone mentions a name, address, or phone number (to protect privacy). It also cuts out long silences so the computer doesn't get bored.
  2. Taking "Voice Photos": The system breaks the call into tiny chunks (like taking a series of snapshots). It turns each chunk of speech into a mathematical "voice photo" (called an embedding). This photo captures the unique shape of the person's voice, ignoring what they are actually saying.
  3. The Grouping Game: Once it has these voice photos from many different calls, it tries to group them.
    • Scenario A (Honest): If Mrs. Smith calls three times, the system groups all three photos together.
    • Scenario B (Fraud): If a fraudster calls as "Mr. Jones" on Monday and "Ms. Brown" on Tuesday, the system notices the voice photos look 99% identical. It groups them together, flagging that the same voice is pretending to be two different people.

How They Tested It

The team didn't just guess; they tested this on 121 real calls from a real insurance company.

  • They had two human experts listen to the calls and group them by who they thought was speaking.
  • They compared the computer's groups to the human experts' groups.
  • The Result: The computer was incredibly accurate. It matched the human experts' grouping about 96% of the time. It was so good at spotting that a voice was "complete" (all parts of the same person) that it got a perfect score of 100% on that specific metric.

What This Means for Investigators

The paper makes it very clear: DGVoiC is not a robot that arrests people.

Instead, think of it as a smart highlighter for human investigators.

  • When an investigator has hundreds of calls to review, DGVoiC scans them and says, "Hey, look! These three calls from different customers all sound like the same person. You should investigate this."
  • It helps investigators spot patterns that are too fast or too subtle for a human to catch while listening to hours of audio.

The Bottom Line

The paper claims that by using this voice-clustering method, insurance companies can better spot when the same voice is being used to pretend to be multiple different people. It works well even with real-world noise and privacy filters. It's a tool to help humans do their job better, not a tool to make the final decision on its own.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →