← Latest papers
💬 NLP

HK-LegiCoST: Leveraging Non-Verbatim Transcripts for Speech Translation

This paper introduces HK-LegiCoST, a large-scale three-way parallel corpus of Cantonese audio, non-verbatim traditional Chinese transcripts, and English translations designed to advance speech translation research for languages where spoken and written forms significantly diverge.

Original authors: Cihan Xiao, Henry Li Xinyuan, Jinyi Yang, Dongji Gao, Matthew Wiesner, Kevin Duh, Sanjeev Khudanpur

Published 2026-06-16
📖 4 min read☕ Coffee break read

Original authors: Cihan Xiao, Henry Li Xinyuan, Jinyi Yang, Dongji Gao, Matthew Wiesner, Kevin Duh, Sanjeev Khudanpur

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot to understand a very specific type of conversation: the formal debates of the Hong Kong Legislative Council. But there's a catch. The people speaking are using Cantonese, a rich, spoken dialect with its own unique slang and sentence structures. However, the official written records of these meetings aren't written in "pure" Cantonese. Instead, they are translated into Standard Chinese, a formal written language that often sounds like a different dialect entirely.

It's like listening to a friend tell a story in a heavy, local accent, but the official transcript you have to match it against is written in a stiff, formal style that rearranges the words and swaps out the slang for polite vocabulary.

This is the challenge the researchers at Johns Hopkins University tackled in their new paper, introducing a massive new dataset called HK-LegiCoST.

The Problem: The "Noisy" Transcript

Usually, when we train computers to translate speech, we assume the written text matches the spoken words perfectly. But in the case of Cantonese, the written record is often a "cleaned-up" version of what was actually said.

  • Spoken: "You go first!" (Cantonese style)
  • Written: "You first go." (Standard Chinese style)

This mismatch is like trying to match a jazz improvisation to a sheet of classical music. The notes are similar, but the rhythm and order are different. This makes it very hard for computers to learn how to translate the speech directly.

The Solution: Building a Giant Puzzle

The researchers took over 600 hours of video recordings from the Hong Kong government meetings. They didn't just copy-paste the files; they built a complex "pipeline" (a step-by-step assembly line) to turn this messy raw data into a clean, usable training set.

Think of their process like a high-tech librarian organizing a chaotic library:

  1. Cutting the Tape: They sliced the long meeting videos into bite-sized audio clips.
  2. The Translator's Match: They used advanced AI to match the audio to the written text. Because the text wasn't a perfect word-for-word match, they had to teach the AI to be flexible. It's like teaching a translator to ignore the fact that the speaker said "y'all" and the text says "you all," and instead focus on the meaning.
  3. The "Skip" Button: Sometimes, the written transcript included notes or formalities that were never actually spoken. The researchers built a special algorithm that acts like a "skip button," allowing the computer to ignore those extra written words and focus only on the parts that were actually said.

The Result: A New Training Ground

The final product, HK-LegiCoST, is a massive library of 600+ hours of Cantonese audio paired with its "imperfect" written transcript and an English translation.

The researchers tested their computer models on this new data and found some exciting things:

  • It Works: Even though the transcripts were "noisy" (not a perfect match), the computer models learned to translate the speech very well.
  • It's Robust: When they tested their model on a different, smaller dataset of Cantonese speech (from Google's FLEURS project), it performed just as well as models trained on much larger, more expensive data.
  • The "Zero-Shot" Magic: A model trained only on this new Hong Kong data was able to handle the Google dataset without any extra training, proving that the data is high-quality and generalizable.

Why This Matters

This paper doesn't just give us a new dataset; it proves that we can build powerful speech tools even when the written records don't perfectly match the spoken words. It's a blueprint for handling languages where the "spoken" and "written" versions are often at odds, which is common in many dialects and vernaculars around the world.

In short, the researchers took a messy, real-world problem (government meetings where the script doesn't match the speech) and turned it into a clean, powerful tool that helps computers understand the true voice of the people, not just the formal words on the page.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →