From Cascaded Speech Systems to Omni-Modal Agents: A Survey of Speech Interaction for Next-Generation Human-Computer Interaction
This survey traces the evolution of speech interaction from traditional cascaded systems to unified omni-modal agents, systematically analyzing the architectures, training strategies, and data governance of SpeechLLMs and Omni-MLLMs while identifying key challenges and future directions for next-generation human-computer interaction.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine trying to have a conversation with a machine that can only hear your words, not your tone, and can only answer in a flat, robotic voice, never interrupting you even if you change your mind mid-sentence. For years, this has been the reality of talking to computers. The technology behind these interactions has traditionally worked like a relay race with three separate runners. First, a system listens to your voice and converts it into written text. Second, a separate computer brain reads that text, figures out what you mean, and writes a reply. Third, a different machine takes that written reply and speaks it back to you. While this method works, it is slow and clumsy. By turning your voice into text and then back into voice, the system loses the subtle cues that make human conversation feel natural: the hesitation in your voice, the rise and fall of your pitch, the emotion behind your words, and the ability to jump in when the other person is still talking.
Now, a new generation of technology is attempting to replace this relay race with a single, unified mind. Researchers are building systems that can listen, understand, and speak all at once, without needing to convert your voice into text first. These systems are designed to perceive the world through multiple senses simultaneously, processing not just your voice, but also images, videos, and environmental sounds to understand the full context of a situation. A new survey of this rapidly evolving field, published by researchers from the City University of Macau and Minzu University of China, maps out how we got here and where we are going. The authors, Sisi Zhu, Yue Zhao, and their colleagues, have gathered the latest research to show how speech interaction is shifting from a rigid, step-by-step process to a fluid, continuous conversation that feels more like talking to a person than operating a machine.
The paper traces the history of this technology through four distinct stages. It began with the traditional "cascaded" systems mentioned earlier, where the three separate modules passed information down the line. The researchers explain that while this approach was easy to build, it suffered from a fundamental flaw: every time the system converted voice to text, it lost some of the richness of the original sound. If the first runner in the relay made a mistake, the error would carry through to the end. The next stage introduced "Speech Large Language Models," which began to integrate the voice directly into the computer's brain, allowing it to understand speech without always turning it into text first. This was a significant leap, preserving more of the speaker's emotion and style.
The field then moved into "Multimodal Large Language Models," which added the ability to see and understand images and videos alongside speech. However, even these systems often treated different types of information separately, struggling to combine them seamlessly. The final and most advanced stage described in the survey is the "Omni-Modal Agent." These are the systems the researchers are most excited about. They are designed to perceive text, images, video, speech, and environmental sounds all at once, within a single framework. Instead of processing one thing at a time, an Omni-Modal agent can listen to you while looking at a video you are showing it, understanding how the sound of your voice matches the scene on the screen. The survey highlights that these systems are beginning to support "full-duplex" interaction, a technical term for the ability to talk and listen at the same time, just like humans do. This means the computer can stop what it is saying if you interrupt it, or it can listen to you while it is still formulating its answer, creating a much more natural flow of conversation.
The researchers did not just describe these technologies; they analyzed how they are built and trained. They found that creating these advanced systems requires massive amounts of data that mix different types of information together. The training process is complex, starting with teaching the model to understand one type of data, like text, and then gradually introducing images, audio, and video. The goal is to teach the model to see the connections between them, such as how a specific sound relates to a visual event. The survey notes that while these models are becoming incredibly capable, they still face significant hurdles. One major challenge is timing. In a real conversation, everything happens in a split second. The researchers point out that aligning the timing of a spoken word with a visual action in a video is difficult, and getting the system to respond instantly without a noticeable delay is still a work in progress.
Another critical finding in the paper concerns the reliability and safety of these new agents. Because these systems are so powerful, they can sometimes "hallucinate," or make up facts, especially when they are trying to combine information from different sources. For example, if a model sees a picture of a dog and hears a sound that might be a cat, it might confidently describe a cat in the picture even if it isn't there. The authors emphasize that building trust in these systems is a top priority. They argue that future models need better ways to verify that their answers are grounded in the actual evidence they see and hear, rather than just guessing based on patterns they have learned. The survey also touches on the issue of privacy, noting that because these systems are listening and watching constantly, they collect a vast amount of sensitive personal data, which raises important questions about security and user control.
Looking toward the future, the researchers suggest that the next big step is not just making these systems smarter, but making them more human-like in their behavior. They envision a future where speech interaction is not just about giving commands, but about having a genuine, ongoing dialogue where the computer understands your mood, remembers your past conversations, and can help you with complex tasks in the real world. The survey concludes that while we are moving away from the old, clunky relay-race style of talking to machines, the journey is far from over. The technology is evolving rapidly, but to truly achieve a natural, trustworthy, and safe interaction, researchers still need to solve problems related to speed, accuracy, and the ethical use of personal data. The path forward involves building systems that are not only powerful but also reliable, ensuring that as machines become better at understanding us, they remain helpful and safe partners in our daily lives.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.