← Latest papers
💬 NLP

DuplexWorld: Can voice agents help you get through the day?

The paper introduces DuplexWorld, a comprehensive benchmark featuring six real-world domains and 156 scenarios to holistically evaluate speech-to-speech voice agents on conversational, analytical, and naturalness metrics, revealing significant performance gaps in current systems despite their growing adoption.

Original authors: Aryan Vijay Bhosale, Harshit Rajgarhia, Akhil Pothanapalli, Asif Shaik, Abhishek Mukherji, Dinesh Manocha

Published 2026-08-12
📖 3 min read☕ Coffee break read

Original authors: Aryan Vijay Bhosale, Harshit Rajgarhia, Akhil Pothanapalli, Asif Shaik, Abhishek Mukherji, Dinesh Manocha

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot how to talk to humans. For a long time, scientists tested these robots by asking them to read text back and forth, like a very fast game of email. But real life isn't text; it's messy, loud, and happens in real-time. This is the world of Speech-to-Speech (S2S) voice agents. Think of them as digital companions that can hear you, understand your voice, and talk back instantly, just like a human on the phone. The big question scientists are asking is: Can these robots actually do things for us, like fix a bank account or guide us through a city, or are they just really good at pretending to listen? We care because if we let these agents handle our daily chores, we need to know if they will actually get the job done without getting lost, forgetting our names, or talking over us.

Enter DUPLEXWORLD, a new "training ground" created by researchers to test these voice agents in the real world. Instead of just asking them to solve math problems on a screen, the researchers built six different "worlds" that mimic our daily lives: Banking, Insurance, Travel, Healthcare, Logistics, and a tricky new one called Pathfinding (where the agent has to guide a walker through a city). They put five different commercial voice agents into these worlds and gave them 156 different scenarios to solve, totaling over 3,800 conversations. It's like a massive, high-stakes video game where the agents have to juggle listening, thinking, and acting all at once.

The results? Even the "best" agents are still a work in progress. The researchers found that being good at conversation doesn't mean you're good at the task. One agent might sound incredibly natural and keep the conversation flowing perfectly, but it might fail to actually book the flight or find the right bank account. In fact, the top-performing agent only succeeded in about 49% of the tasks on its first try (a score of 0.490). Another key finding is that "sounding good" is a trap; the agents with the highest voice quality scores weren't necessarily the ones that completed the tasks best. In the Pathfinding world, the agents that spent too much time "exploring" or asking too many questions actually got lost more often, while the ones that stuck to a plan arrived at their destination.

The paper also explicitly rules out the idea that we can just judge these agents by how polite or smooth they sound. The data shows that an agent can be a "chatterbox" that never actually does anything, and if we only looked at conversation scores, we might pick the wrong robot for the job. Furthermore, the researchers found that even the best systems struggle with reliability; if you ask them to try the same task three times, they often fail at least once. The study suggests that for voice agents to truly help us, they need to get much better at the actual work, not just the talking. It's a reminder that while the technology is impressive, we aren't quite ready to hand over our daily lives to them just yet.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →