UESF-Bench: Benchmarking and Probing for Unified Embodied Seeking and Following
This paper introduces UESF-Bench, a large-scale benchmark for unified embodied human seeking and following that addresses the limitations of existing evaluations by requiring agents to first locate a language-described target before persistently tracking them, and proposes the SeekFollow-VLA framework to effectively manage the transition between these two phases.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are teaching a robot to be the ultimate sidekick. In the world of robotics, there's a big difference between just "seeing" and "understanding." For a robot to be truly helpful, it needs Embodied AI—a brain that lives inside a body, allowing it to move through the real world, not just stare at a screen. It also needs Vision-Language-Action skills, which is a fancy way of saying it must look at what's happening, listen to what you say, and then physically move to do something about it. Think of it like a game of hide-and-seek where the robot has to listen to a clue like "Find the guy in the red jacket," run around a house to find him, and then stick close by without bumping into him. Why does this matter? Because in real life, people don't always stand right in front of a robot waiting to be followed. Often, they are hiding in another room, or the robot has lost sight of them. If a robot can only follow someone it can already see, it's not very useful. It needs to be able to hunt them down first, then switch gears to follow them, all while figuring out who is who in a crowded room.
This paper introduces a new challenge and a new solution for exactly that problem. The authors, a team of researchers from various universities and companies, realized that most existing tests for robot followers were too easy: they assumed the target person was already standing right there, visible from the start. They argue that this misses the most important part of the job: the search. To fix this, they built a massive new playground called UESF-Bench (Unified Embodied Seeking and Following Benchmark). It's like a giant video game level with over 1.43 million different scenarios, featuring 4,800+ unique digital people and 770+ different environments. In this benchmark, the robot gets a single instruction like, "Find the short man with brown hair in the kitchen and follow him," and it has to do two things: first, actively search for him when he's out of sight, and second, seamlessly switch to following him once found, even if he gets lost in a crowd again.
The researchers found that simply telling a robot to "do both" isn't enough; it gets confused. They tested three different ways to build the robot's brain. The first was a "Single Head," where the robot tried to use one brain for both searching and following. The second was a "Dual Head" with a basic switch, and the third was a "Task-Driven Router," a smarter system that acts like a traffic cop, explicitly telling the robot when to stop searching and start following based on what it sees. The results showed that the "Single Head" approach was a disaster, with the robot succeeding in only 4% of the single-person tests. Even the basic "Dual Head" struggled, staying at 5%. However, the "Task-Driven Router" (which they call SeekFollow-VLA) was a game-changer. In single-person tests, it succeeded 35% of the time, and in the much harder multi-person tests (where fake people are hiding to trick the robot), it still managed a 20% success rate, far beating the others.
The paper suggests that the key to success isn't just having a better camera or a faster brain, but having a clear internal signal that knows when the job has changed from "hunting" to "herding." The researchers showed that their smart router learns to act like a detective first, scanning the room with high intensity, and then instantly switches to a loyal bodyguard mode the moment it spots the right person. They also noted that while this new method is much better at finding and following, it does bump into things slightly more often in crowded rooms, suggesting that being a great hunter sometimes makes you a bit clumsy. Ultimately, the paper argues that to build robots that can truly help us in the real world, we need to stop testing them on easy, static tasks and start challenging them with the messy, confusing reality of finding someone who is hiding, and then sticking with them through the chaos.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.