← Latest papers
💬 NLP

Poster: Exploring the Limits of Audio-Based Detection of Turkish Phone Call Scams

This paper introduces the first public multi-modal dataset of Turkish scam calls and evaluates seven large language models, finding that transcript-based inputs consistently outperform raw audio processing for detecting phone scams in low-resource languages.

Original authors: Arda Eren, Micheal Cheung, Youqian Zhang, Grace Ngai, Eugene Yujun Fu

Published 2026-06-24
📖 4 min read☕ Coffee break read

Original authors: Arda Eren, Micheal Cheung, Youqian Zhang, Grace Ngai, Eugene Yujun Fu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine phone scams as a global game of "Wolf in Sheep's Clothing." Scammers use tricks, fake voices, and urgent stories to trick people out of their money. For a long time, researchers have built "wolf detectors" (AI systems) to catch these calls, but they've mostly only learned to spot wolves speaking English or other common languages.

This paper is about building a detector for Turkish, a language that hasn't gotten much attention in the AI world yet. Here is the story of what they did and what they found, explained simply:

1. The Missing Puzzle Piece

The researchers realized that to catch Turkish scammers, they needed a special training manual. Since none existed, they created the first-ever public collection of 100 Turkish phone calls.

  • The Collection: It's a mix of 50 real scam calls and 50 harmless calls.
  • The Source: They grabbed these from YouTube videos where people had already recorded the scams.
  • The Check: A native Turkish speaker listened to every single one to make sure the "scam" ones were actually scams and the "harmless" ones were safe.

2. The Three Ways to Listen

To see how well modern AI (called Large Language Models or LLMs) could spot these scammers, they tested the AI in three different ways, like trying to identify a thief by:

  1. Listening to the raw voice: Feeding the actual audio file directly to the AI.
  2. Reading a rough draft: Using a robot to type out what was said (Speech-to-Text) and giving that messy text to the AI.
  3. Reading a polished report: Having a human fix the robot's typos and then giving that perfect text to the AI.

They tested this against seven different AI models (like different brands of smart assistants) to see which method worked best.

3. The Big Surprise: Text Wins, Audio Loses

The results were a bit counter-intuitive. You might think listening to the tone of a voice (is he shouting? is he nervous?) would be the best way to catch a liar. But the paper found the opposite:

  • The Text Champions: When the AI read the words (either the rough robot draft or the human-corrected version), it was almost perfect. It caught the scammers with a success rate of nearly 99%.
  • The Audio Struggles: When the AI listened to the raw audio, its performance dropped slightly (to about 97%).

Why did the audio fail?
The paper suggests two main reasons, using a "Safety Guard" analogy:

  • The Over-Protective Bouncer: Real scammers often use scary tactics, like shouting, swearing, or pretending to be police to intimidate victims. When the AI heard these aggressive sounds, its internal "safety guard" got scared and refused to listen to the call at all. It treated the scam call like a dangerous threat and shut down.
  • The Silent Text: When the same scary words were written down as text, the "safety guard" didn't panic. It just read the words and correctly identified, "Oh, this is a scam."
  • The Noise Factor: Real phone calls have background noise and people talking over each other. The AI sometimes got confused by the audio static, whereas text is clean and clear.

4. The Human Touch Didn't Matter Much

The researchers wondered if they needed a human to fix the robot's typos (Method 3) to get good results. The answer was no.

  • The AI did just as well with the rough robot draft as it did with the human-corrected version.
  • This means that for catching scammers, you don't need to pay a human to clean up the text first; the AI is smart enough to handle the messy robot version.

5. The Bottom Line

This study shows that for Turkish phone scams, reading the transcript is a better strategy than listening to the audio.

The main takeaway isn't that audio is useless, but that current AI safety systems are too sensitive to the sound of aggression (shouting, swearing) that scammers use. These safety systems accidentally block the very calls they need to analyze. The paper concludes that to protect people in languages like Turkish, we need AI that can handle these real-world, messy, and sometimes aggressive conversations without getting scared and shutting down.

In short: To catch a Turkish scammer, it's currently better to read the script of their lies than to listen to their scary voice, because the AI's "safety filters" get too jumpy when they hear the scammers shouting.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →