← Latest papers
💬 NLP

Syn-TurnTurk: A Synthetic Dataset for Turn-Taking Prediction in Turkish Dialogues

This paper introduces Syn-TurnTurk, a synthetic Turkish dialogue dataset generated by Qwen LLMs to address the scarcity of resources for turn-taking prediction, demonstrating that advanced models trained on this data achieve high accuracy in managing natural dialogue timing and reducing interruptions.

Original authors: Ahmet Tuğrul Bayrak, Mustafa Sertaç Türkel, Fatma Nur Korkmaz

Published 2026-04-16
📖 4 min read☕ Coffee break read

Original authors: Ahmet Tuğrul Bayrak, Mustafa Sertaç Türkel, Fatma Nur Korkmaz

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are having a conversation with a robot. You know that awkward moment when you're still finishing your sentence, but the robot thinks you're done and jumps in to talk? It feels like being interrupted by a friend who can't wait to tell their own story. That's the problem this paper tries to solve, but specifically for Turkish speakers.

Here is the story of Syn-TurnTurk, explained simply:

1. The Problem: The Robot That Can't "Read the Room"

Most voice assistants today are like nervous students waiting for the teacher to say "Stop." They listen for silence. If you stop talking for even a split second, they think, "Okay, silence means I'm done!" and they start talking.

But humans are messy. We pause to think, we say "um," we overlap with each other, and we sometimes stop mid-sentence to add a thought. Because of this, robots often interrupt us, making the conversation feel robotic and annoying.

The real issue? There are plenty of data sets for English to teach robots how to handle these pauses, but Turkish is like a lonely island in the world of data. There just aren't enough examples of Turkish people talking to each other to teach the robots the specific "rhythm" of the language.

2. The Solution: Building a "Virtual Playground"

Since they couldn't find enough real conversations, the authors decided to build their own.

They used powerful AI brains (called Qwen Large Language Models) to act as actors. They gave these AI actors a script with 79 different topics (like "cooking," "travel," or "politics") and told them:

  • "Don't just talk in perfect, robotic lines."
  • "Interrupt each other sometimes."
  • "Pause in weird places."
  • "Use filler words like 'uh' and 'ah'."

They created a massive library of 1,625 fake conversations (called Syn-TurnTurk) that sound surprisingly real. It's like a rehearsal room where thousands of virtual Turkish people practiced talking over each other, pausing, and finishing sentences, just so a robot could learn the rules of the game.

3. The Test: Teaching the Robot to Dance

Once they had this "playground" of fake conversations, they needed to teach a computer model how to predict when a turn ends.

Think of it like a dance partner. You don't want to step on your partner's toes (interrupt), but you also don't want to stand there awkwardly waiting for them to finish if they've already stopped.

They tested different "brains" (algorithms) on this new data:

  • Simple Brains: Like a basic decision tree (a flowchart). These were okay but often got confused.
  • Deep Learning Brains: Specifically BI-LSTM and Ensemble models. These are like experienced dancers who can feel the rhythm, the speed, and the subtle cues in the language.

4. The Results: A New Rhythm

The results were promising. The advanced models (especially the BI-LSTM) got about 84% accuracy.

What does that mean? It means the robot is finally learning to listen to words and grammar, not just silence.

  • In Turkish, the end of a sentence often has specific grammatical markers (like suffixes). The new models learned to spot these clues.
  • They learned that a 0.5-second pause might just be someone thinking, but a 2-second pause with a specific word ending means, "I'm done, your turn!"

The Big Takeaway

This paper is like building a training simulator for Turkish voice assistants. By creating a synthetic dataset that mimics the messy, overlapping, and pausing nature of real human speech, the authors gave robots a chance to learn the "dance" of conversation.

Instead of just waiting for silence, the next generation of Turkish chatbots might finally learn to listen to the music of the conversation, knowing exactly when to step in and when to let the other person finish their sentence.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →