CallBench: A Benchmark for Dual-Goal Coordination in Phone Call Assistants
This paper introduces CallBench, a comprehensive Chinese benchmark comprising 50,000 multi-turn dialogues across six scenarios to evaluate the challenging dual-goal coordination capabilities of phone call assistants, revealing that current methods struggle to effectively balance the device owner's explicit preset goals with the caller's implicit and dynamic objectives.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are the conductor of a very tricky orchestra. In the world of computer science, this is called building a "dialogue system"—a robot that can talk to humans. For a long time, scientists have taught these robots how to listen to a single person and help them finish a task, like booking a flight or ordering a pizza. It's like a robot playing a duet with one musician; they know the song, they know the goal, and they just need to hit the right notes. But what happens when the robot isn't just talking to one person? What if it's standing in the middle of a busy street, holding a phone for someone else, trying to listen to a stranger on the other end of the line while also remembering a secret list of instructions from the person who owns the phone? This is the messy, real-world challenge of a "proxy" setting. It's not just about finishing a task; it's about juggling two different people's needs at once without saying the wrong thing, revealing secrets, or making promises you can't keep.
This is exactly the problem a new study tackles. The researchers introduce a massive new testing ground called CALLBENCH, designed specifically to see how well AI assistants can handle these double-duty phone calls. They created 50,000 complete, multi-turn phone conversations in Chinese, covering six different scenarios like ordering food, catching a taxi, dealing with work calls, and even handling harassment. The goal was to see if current AI models could figure out when to stick to the phone owner's preset instructions (like "leave the package at the door") and when to pause those instructions because the person calling has a problem (like "the package is broken").
The study suggests that while current AI models are getting better at talking, they are still struggling with this specific "dual-goal" juggling act. The researchers found that existing methods often fail to balance the two goals, sometimes blindly pushing the owner's instructions even when the caller is in distress, or conversely, getting so distracted by the caller that they forget the owner's instructions entirely. They also discovered that safety is a huge hurdle; the AI often accidentally crosses the line into making unauthorized decisions or revealing private info.
To test this, the team didn't just look at whether the call ended successfully. They built a detailed scoring system that judges the AI on seven different levels: did it understand what was said? Did it remember the context? Did it guide the conversation at the right time? Was the response natural? Did it follow the owner's rules? Did it keep the conversation moving at a good pace? And most importantly, was it safe?
The results were a bit of a wake-up call. Even the smartest AI models tested, including some that are usually very good at planning and reasoning, scored surprisingly low on the overall test. The most common mistakes weren't just about sounding robotic; they were about bad timing and safety. For instance, the AI would often repeat the owner's preset message even when the caller had just reported an emergency, or it would try to end the call too quickly before all the details were sorted out. The study suggests that to build a truly reliable phone assistant, we need models that can make careful, turn-by-turn decisions about which goal to prioritize, all while staying strictly within the safety boundaries of a "proxy" who is just a messenger, not a decision-maker. It turns out that being a good phone assistant is less about having a big vocabulary and more about knowing exactly when to speak, when to listen, and when to stay silent.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.