← Latest papers
💬 NLP

AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use

This paper introduces AppWorld-UL, a challenging benchmark featuring 516 tasks across nine simulated applications that evaluates the ability of tool-use agents to handle diverse, necessary interactions with users, revealing that even state-of-the-art models struggle significantly with these user-in-the-loop scenarios.

Original authors: Junzhi Chen, Harsh Trivedi, Jane Pan, Michael JQ Zhang, Tejas Srinivasan, Niranjan Balasubramanian, Ashish Sabharwal

Published 2026-07-24
📖 4 min read☕ Coffee break read

Original authors: Junzhi Chen, Harsh Trivedi, Jane Pan, Michael JQ Zhang, Tejas Srinivasan, Niranjan Balasubramanian, Ashish Sabharwal

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a super-smart robot butler how to run your errands. You tell it, "Go buy me a shirt." In the world of computer science, this is called "tool use"—the robot knows how to open an app, click buttons, and make a purchase. But here's the catch: real life is messy. What if you have five shirts in your cart and the robot doesn't know which one you mean? What if the shirt you want is sold out? What if the price suddenly doubles, and you'd only buy it if you say "yes" to that specific change?

Most tests for these robots only check if they can follow a perfect, step-by-step recipe where nothing goes wrong. But real humans are unpredictable. We get vague, we change our minds, and we need to be asked for permission before the robot spends our money. This paper steps into that messy, real-world corner of AI research to ask a simple but tough question: Can our best AI agents actually talk to us when things get complicated, or do they just crash and burn when the instructions aren't perfect?

The researchers behind this study, led by Junzhi Chen and Harsh Trivedi, decided to stop testing robots on perfect, imaginary days and start testing them on days that feel like real life. They built a new playground called AppWorld-UL (User-in-the-Loop). Think of it as a giant, simulated digital city with nine popular apps like Amazon and Spotify, but with a twist: the "user" (a simulated person) is part of the game, and the robot has to figure out how to ask them for help.

The team didn't just make up random problems; they used a clever "perturbation" method. Imagine they took a normal task, like "Return the Nike shirt I bought last week," and then secretly tweaked the world to make it tricky.

  1. The "Which One?" Trap (Underspecification): They made sure the robot found two Nike shirts from last week. Now the robot can't just guess; it has to stop and ask, "Which one do you want to return?"
  2. The "It's Gone!" Trap (Infeasibility): They made the "Large" size of the shirt disappear from the store. The robot can't just fail; it has to tell the user, "The large size is gone. Do you want a different size or a refund?"
  3. The "Wait, That's Expensive!" Trap (Confirmation): They made the price of the shirt double. The robot has to pause and say, "Hey, the price jumped! Do you still want to buy it?"

To make sure the test was fair, they created a "simulated user" that acts like a real person but follows strict rules. This user knows the answers to the tricky questions but won't give them away unless the robot asks the right question. It's like a game of "20 Questions" where the user will only answer if you ask exactly what they are thinking about.

When they put the world's smartest AI robots (like Claude Opus 4.7 and GPT-5.5) into this chaotic digital city, the results were a bit of a wake-up call. Even the best robot only succeeded about 48.6% of the time on the whole set of tasks. When the tasks got even harder, combining two or three of these tricky situations at once, the success rate dropped to just 35.7%.

The study suggests that the problem isn't that the robots can't use the apps; they are actually pretty good at clicking buttons. The real struggle is the conversation. When the researchers gave the robots a "cheat sheet" with all the missing information upfront (so they didn't have to ask the user), the success rate jumped to 78.1%. This proves that the difficulty comes from the need to interact, not from the complexity of the apps themselves.

The paper also found that when robots failed, it was usually because they didn't ask the right questions. They either guessed the answer (hallucinating), asked vague questions that confused the user, or forgot what the user just told them. It's like a robot that hears "I want the red shirt" but then buys the blue one because it forgot to check.

In short, this paper shows that while our AI agents are getting great at doing tasks, they are still learning how to be good conversational partners. They need to learn that sometimes the best way to solve a problem isn't to rush ahead, but to stop, look at the human, and ask, "Wait, which one did you mean?" Until they master that back-and-forth dance, they might be ready for a simple grocery run, but they aren't quite ready to handle the chaos of a real human's day.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →