← Latest papers
⚡ electrical engineering

TELEVAL: A Benchmark Designed for Spoken Language Models in Chinese Interactive Scenarios

This paper introduces TELEVAL, a large-scale Chinese benchmark designed to evaluate spoken language models on both semantic accuracy and interactional appropriateness in audio-conditioned settings, revealing that current models struggle with acoustic variability and often fall into a "Caption Trap" where they describe audio rather than engaging naturally.

Original authors: Zehan Li, Hongjie Chen, Qing Wang, Yuxin Zhang, Jing Zhou, Hang Lv, Mengjie Du, Yaodong Song, Jie Lian, Jian Kang, Jie Li, Yongxiang Li

Published 2026-08-25
📖 5 min read🧠 Deep dive

Original authors: Zehan Li, Hongjie Chen, Qing Wang, Yuxin Zhang, Jing Zhou, Hang Lv, Mengjie Du, Yaodong Song, Jie Lian, Jian Kang, Jie Li, Yongxiang Li

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine sitting in a bustling teahouse in Chengdu. A friend leans in, speaking in a thick local dialect, their voice slightly raspy from a cold, and they sigh with a hint of frustration before asking a simple question. A human listener would instantly catch the dialect, the tone, the cough, and the sigh, weaving all those clues into a warm, natural reply. They wouldn't just hear the words; they would hear the person. For years, computers have been getting better at understanding the words we speak, but they often miss the rest of the conversation. They can tell you the answer to a question, but they struggle to know how to say it, or when to offer comfort, or how to adjust their voice to match the mood of the room. This gap between simply hearing words and truly understanding a spoken interaction is the central challenge for the next generation of artificial intelligence.

A team of researchers from China Telecom has built a new tool to measure exactly how well computers are doing at this difficult task. They call it TELEVAL. Unlike previous tests that focused on whether a machine could correctly identify a sound or answer a multiple-choice question, this new benchmark asks a different question: does the machine act like a natural conversational partner? The researchers created over 40,000 examples of spoken interactions, ranging from a child asking for help to an elderly person discussing health insurance, all recorded in various Chinese dialects and emotional states. They tested a dozen different spoken language models, including some of the most advanced systems available today, to see how they handled these real-world scenarios.

The results revealed a stark divide. While these artificial intelligence models performed quite well when asked to simply identify facts or answer questions in a quiet, controlled setting, their performance crumbled when the situation became messy and human. When the audio contained background noise, when the speaker used a regional dialect like Sichuanese or Shanghainese, or when the user's voice carried subtle cues like a cough or a sigh, the models often failed to adapt. They would give a technically correct answer but miss the point entirely, or they would respond with a robotic tone that ignored the user's distress. The study found that the more complex the interaction became, the more the machines struggled to bridge the gap between hearing and understanding.

One of the most striking discoveries was a specific failure pattern the researchers named the "Caption Trap." In many instances, when a model heard a sound like a sneeze or a cough, it would stop and describe the sound to the user, saying something like, "I hear you sneezing," rather than responding with care, such as, "Are you okay?" or "Take a moment to rest." It was as if the machine was acting like a reporter describing a scene rather than a friend having a conversation. This behavior suggested that the models were trained to recognize and label sounds, but they had not learned how to use those sounds to guide their social behavior. They could perceive the signal, but they could not translate that perception into a helpful action.

The study also highlighted how sensitive these systems are to the way people speak. When the input switched from standard Mandarin to a local dialect, the accuracy of the models dropped significantly. Some systems performed reasonably well with dialects that are close to standard speech, but they faltered completely with others, such as Shanghainese. Similarly, when the models were asked to adjust their tone for a child versus an adult, they often failed to change their vocabulary or style, sticking to a generic adult voice even when speaking to a child. This suggests that current technology is still heavily biased toward standard, clear speech and has not yet learned to navigate the rich diversity of human communication.

Perhaps most importantly, the research showed that a model's ability to answer a question correctly does not guarantee it can hold a conversation. The rankings of the models changed drastically depending on the type of test. A system that ranked highly on traditional benchmarks for knowledge and reasoning often fell to the bottom when tested on its ability to be empathetic or natural. This indicates that the skills required to pass a test are different from the skills required to be a good conversational partner. The models are currently optimized to be accurate, but they are not yet optimized to be human-like.

The researchers conclude that while spoken language models have made impressive strides in understanding the content of speech, they remain insufficiently aligned with the requirements of natural interaction. They can hear the words, but they are still learning how to listen to the person. By exposing these specific weaknesses, from the inability to handle background noise to the tendency to describe sounds instead of reacting to them, the TELEVAL benchmark provides a clear roadmap for future improvements. It suggests that the next step in artificial intelligence is not just about knowing more facts, but about learning to respond with the same nuance, care, and adaptability that humans use every day in their conversations.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →