← Latest papers
🤖 machine learning

PhoneWorld: Scaling Phone-Use Agent Environments

PhoneWorld introduces a scalable pipeline that automatically converts real mobile GUI trajectories into controllable, verifiable training environments across 34 apps, significantly improving agent performance on multiple benchmarks by shifting the focus from manual benchmark construction to mass-producing phone-use environments.

Original authors: Zhengyang Tang, Yuxuan Liu, Xin Lai, Junyi Li, Pengyuan Lyu, Jason, Yiduo Guo, Zhengyao Fang, Yang Ding, Yi Zhang, Weinong Wang, Huawen Shen, Xingran Zhou, Liang Wu, Fei Tang, Sunqi Fan, Shangpin Pen
Published 2026-05-29
📖 4 min read☕ Coffee break read

Original authors: Zhengyang Tang, Yuxuan Liu, Xin Lai, Junyi Li, Pengyuan Lyu, Jason, Yiduo Guo, Zhengyao Fang, Yang Ding, Yi Zhang, Weinong Wang, Huawen Shen, Xingran Zhou, Liang Wu, Fei Tang, Sunqi Fan, Shangpin Peng, Zheng Ruan, Anran Zhang, Benyou Wang, Rui Yan, Ji-Rong Wen, Chengquan Zhang, Han Hu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you want to teach a robot how to use a smartphone. The robot needs to see the screen, understand what buttons do, and figure out how to get things done—like booking a flight, ordering food, or sending a message.

The big problem, according to this paper, isn't just that the robot is "dumb." It's that we don't have enough practice fields for it to learn in. Real phone apps are messy. They change all the time, they are hard to reset to a "fresh start" state, and they are expensive to turn into a safe training ground. It's like trying to teach someone to drive only by letting them practice on real, busy highways with no way to pause or rewind if they crash.

PhoneWorld is the solution the authors built. Think of it not as a single test, but as a factory that builds practice fields.

Here is how it works, using simple analogies:

1. The Blueprint (Learning from Real Life)

Instead of guessing what a phone app looks like, the team watched real people use real apps. They recorded the screens people looked at and the paths they took (like going from the home screen to a search bar, then to a product page).

  • The Analogy: Imagine you want to build a perfect replica of a popular restaurant for a cooking class. Instead of guessing the menu, you watch hundreds of real customers, see which tables they sit at, what they order, and how the waiter moves between the kitchen and the dining room. PhoneWorld does this with apps.

2. Building the "Mock" Apps

Using those observations, PhoneWorld automatically builds fake versions of real apps (like a fake QQ, a fake shopping app, or a fake travel app).

  • The Magic: These fake apps look and feel like the real ones, but they are resettable. If the robot makes a mistake, you can hit "reset," and the app goes back to exactly how it started.
  • The Database: Inside these fake apps, there is a hidden "scoreboard" (a database). If the robot successfully sends a message or adds an item to a cart, the scoreboard updates automatically. This lets the computer know instantly, "Yes, the robot did it right," without a human needing to watch.

3. The Training Loop

Because these environments are safe and resettable, the robot can practice thousands of times.

  • The Process: The robot tries to complete a task (e.g., "Find a movie and buy a ticket"). If it succeeds, the system saves that success as a lesson. If it fails, it tries again.
  • The Scale: The team built 34 different fake apps covering 16 different categories (shopping, social media, travel, etc.). This gives the robot a huge variety of "driving lessons" to learn from.

What They Found (The Results)

The team ran experiments to see if this "factory of practice fields" actually made the robot smarter. They compared a robot trained on old data versus one trained with their new PhoneWorld data.

  • The "Mix-and-Match" Test: They took a strong robot and swapped out just a small chunk of its old training data (10,000 steps) with new PhoneWorld practice.
    • The Result: The robot got significantly better at everything. It didn't just get better at the fake apps; it got better at real-world tests too. It was like swapping a few hours of driving on a quiet track for a few hours of driving in a busy city simulation—the robot learned to handle real traffic better.
  • The "More is Better" Test: They found that simply adding more PhoneWorld practice helped, but adding more types of apps helped even more.
    • The Analogy: It's not just about practicing driving for 100 hours on a single straight road. It's about practicing on a highway, a dirt road, a rainy street, and a parking lot. The variety of the environments was the secret sauce.

The Big Takeaway

The paper argues that to make phone-using robots smarter, we shouldn't just build bigger models or collect more static data. Instead, we need to scale the supply of practice environments.

PhoneWorld is a tool that turns real-world usage into a reusable, resettable, and automatic training gym. It allows us to build many different "phone worlds" quickly, giving robots the diverse, safe, and endless practice they need to master the real thing.

In short: PhoneWorld is a factory that builds safe, resettable video game versions of real phone apps, allowing AI robots to practice millions of times so they can eventually use real phones without crashing.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →