← Latest papers
💬 NLP

Modeling turn-taking with distant viewing: investigating silence thresholds in human and AI-generated discourse

This study investigates silence thresholds in human and AI-generated discourse by analyzing and comparing gap durations across thirty US sitcoms and fifty-one synthetic podcasts, examining variations based on speaker gender and production settings.

Original authors: Taylor Arnold, Nicolas Ballier, Artem Saloev

Published 2026-07-21
📖 4 min read☕ Coffee break read

Original authors: Taylor Arnold, Nicolas Ballier, Artem Saloev

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are at a busy party where everyone is chatting. You know the rhythm of a good conversation: someone stops talking, there's a tiny, almost invisible pause, and then someone else jumps in. It happens so fast you barely notice it. Scientists who study how people talk call this "turn-taking." They've spent decades measuring these tiny pauses to understand how humans coordinate their speech without stepping on each other's toes. Usually, they just listen to the audio, like a radio listener trying to figure out who is speaking. But what if the "silence" you hear isn't actually the speakers pausing? What if the silence is actually the director of a movie hitting the "cut" button, or a computer program deciding to switch voices? This is the question a team of researchers asked. They wanted to see if the tiny gaps between speakers in TV shows and AI-generated podcasts are really about the people talking, or if they are just the result of how the audio and video were edited together.

The researchers, Taylor Arnold, Nicolas Ballier, and Artem Saloev, decided to play detective with two very different kinds of audio. First, they looked at thirty classic American sitcoms (like Friends or The Office), which are human-made, scripted, and heavily edited. Second, they analyzed fifty-one synthetic podcasts created by an AI tool called Google's NotebookLM, where computer voices chat with each other. They used a special computer program to listen to the audio and measure the exact length of the silence between one speaker stopping and the next starting. But they didn't stop there. For the TV shows, they also watched the video. They looked for "shot cuts"—the moment the camera switches from one view to another, like when the camera jumps from a person's face to their friend's face.

Here is what they found, and it's a bit of a plot twist. When they looked at the TV shows, they discovered that the "silence" between speakers wasn't just about the actors waiting for their turn. A huge chunk of that silence was actually the editor's doing. When the silence happened right at the exact moment the camera cut to a new shot, the pause was much longer—about 711 milliseconds (a little less than a second) on average. But when the silence happened while the camera was still staring at the same scene (a "mid-shot" gap), it was much shorter, around 424 milliseconds. That's a difference of 287 milliseconds, which is a massive gap in the world of conversation. It turns out that the "rhythm" of a TV show is often set by the editor cutting the film, not by the actors deciding when to speak. The camera cutting adds a "breath" of silence that the computer program mistakenly thinks is the speakers pausing.

The AI podcasts told a different story. The computer-generated conversations had much shorter pauses, with cross-gender transitions (like a female voice switching to a male voice) happening in about 287 to 310 milliseconds. These gaps were uniform and mechanical, lacking the variety found in human shows. The AI didn't have an editor to cut the video, and it didn't have a studio audience to laugh at, so it just switched voices as quickly as possible. It didn't mimic the natural, negotiated timing of humans; instead, it flattened the conversation into a rapid, robotic handoff.

The study suggests that if you only listen to the audio of a TV show, you might get the wrong idea about how people actually talk. You might think the speakers are taking long pauses, when really, the editor just cut the camera. The researchers point out that their study has limits: neither the TV shows nor the AI podcasts had much "overlap" (where people talk at the same time), which is a huge part of real, messy human conversation. But the main takeaway is clear: in edited media, the "silence" you hear is often a product of the camera and the editor, not the speakers. To truly understand how humans take turns talking, we need to look at unscripted, spontaneous conversations where the camera isn't cutting the rhythm for us.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →