Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing
This paper proposes treating transcription style as a controllable latent variable in ASR models, demonstrating that coverage-aware decoder tokens and supervised cross-attention fine-tuning can significantly improve verbatim accuracy, disfluency detection, and word-level timing while introducing a new task to scale the creation of high-quality verbatim speech corpora.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are talking to a very smart, very fast robot that listens to your voice and types out exactly what you say. This robot is part of a field called Automatic Speech Recognition (ASR). For a long time, these robots had a confusing habit: they didn't know whether to type out your words exactly as they came out of your mouth—including your "umms," "uhhs," stutters, and self-corrections—or to clean everything up and just give you the smooth, perfect sentence you meant to say. It was like asking a scribe to write down a story, but you never told them if they should include the author's nervous coughing or just the plot. This confusion made the robot's typing shaky, its timing on when words started and stopped unreliable, and it was hard to tell if the robot was actually making mistakes or just choosing a different style of writing.
The big question researchers have been asking is: Is the robot actually bad at hearing the messy parts of speech, or is it just confused about the rules? This new paper suggests that the robot isn't broken; it's just been playing a game with mixed-up instructions. The authors argue that the ability to hear and type out those messy, real-life speech sounds is already hidden inside the robot's brain, but it needs a specific "switch" to turn it on. They treat the choice between "messy verbatim" and "clean intended" speech not as a missing skill, but as a hidden setting that can be flipped.
The team, working with a popular AI model called Whisper, decided to test this theory by giving the robot a simple control panel. They introduced special "mode tags"—tiny, invisible instructions that tell the robot exactly what to do before it starts typing. If you say "verbatim mode," the robot types everything, including the stutters and pauses. If you say "intended mode," it cleans up the speech and gives you the smooth version. They found that by simply adding these tags, they could make the robot switch between these styles perfectly, even for languages it had never seen before.
Here is the magic part: The researchers discovered that they didn't need to re-teach the robot how to hear. In fact, they froze the robot's brain completely and only trained the tiny "mode tag" switches. With zero changes to the robot's main memory, they managed to boost its ability to detect speech stutters in German from a terrible 10% success rate to a fantastic 79% success rate, just by using English training data. This suggests the skill was there all along, waiting to be activated.
They also fixed the robot's sense of time. Usually, when a robot stumbles over a word, it gets confused about exactly when that word started and ended. By teaching the robot to pay attention to its own internal "focus" while it types, they made its timing incredibly precise. They even created a new trick called "Verbatimize." This allows the robot to take a clean, perfect sentence and a recording of a messy speech, and then reconstruct the messy version from scratch, inserting the stutters and pauses exactly where they happened. This is huge because it means we can turn thousands of clean transcripts into messy, realistic ones without needing humans to listen to every single recording and type them out manually.
The paper shows that the biggest problem wasn't that the robots lacked the ability to handle messy speech, but that we hadn't given them a clear way to choose how to handle it. By making the "transcription policy" an explicit choice rather than a hidden guess, the researchers made the robot more stable, more accurate, and much better at understanding the full, human reality of how we speak.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.