← Latest papers
💻 computer science

EgoAfford: Task-Oriented Affordance Grounding via Egocentric Referring Segmentation

The paper introduces EgoAfford, a benchmark and dataset for task-oriented affordance grounding in egocentric views, along with EgoLens, a specialized multimodal model that jointly performs next-step planning and role-specific segmentation for multi-step tabletop tasks.

Original authors: Xinyuan Guan, Feifan Chen, Xinyu Zhan, Fu-Cheng Zhang, Cewu Lu, Lixin Yang

Published 2026-08-06
📖 3 min read☕ Coffee break read

Original authors: Xinyuan Guan, Feifan Chen, Xinyu Zhan, Fu-Cheng Zhang, Cewu Lu, Lixin Yang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a robot to make a sandwich. You don't just want the robot to know what a "knife" looks like; you need it to understand how to use that knife to spread butter, what to spread it on, and where to put the finished slice. This is the world of affordance, a fancy word for "what an object allows you to do." Think of a doorknob: its shape affords turning. A cup affords holding liquid. For a long time, robots have been pretty good at spotting these simple, one-off actions, like "grab the handle." But real life is messy and multi-step. Making a sandwich isn't just one grab; it's a whole story involving planning, picking up the right tool, and knowing exactly where to put things next. The big question scientists are asking is: Can we teach robots to not just see objects, but to understand the story of a task, figure out the next move, and point out the tiny, specific parts of an object needed for that move?

Enter EgoAfford, a new project that tries to solve this puzzle by giving robots a "first-person" view of the world, just like we see it. The researchers built a massive digital playground with about 15,500 images of tabletop scenes, covering 2,000 different multi-step tasks. They also created a "real-world" test set with 102 photos taken by humans to see if the robot can handle the messy reality of a real kitchen. The goal? To see if a computer can look at a picture, read a goal like "make tea," and then do two things at once: 1) Write down the rest of the recipe (the plan), and 2) Draw a mask over the exact part of the teapot you need to pour from, the spoon you need to stir with, and the cup you need to pour into.

To tackle this, the team introduced EgoLens, a smart AI model designed specifically for this job. Think of EgoLens as a robot chef's apprentice who is learning to "read" a scene. Unlike other models that might guess the whole object or get confused by complex instructions, EgoLens is trained to break a task down into its smallest functional pieces. It doesn't just say "that's a pot"; it says, "I need to grab the spout of that pot to pour, and I need to hold the handle of the spoon to stir." The paper shows that EgoLens is currently the best at this specific type of thinking, outperforming other big AI models and even commercial tools that try to combine planning and vision.

However, the researchers are careful to note that this isn't a magic solution yet. While EgoLens is great at the "digital" images they generated, there is still a gap when it comes to real-world photos, suggesting that the robot still needs more practice with the unpredictable nature of real life. The study suggests that by combining the ability to plan the next steps with the ability to pinpoint the exact functional part of an object, we can get robots much closer to being helpful helpers in our daily lives. It's a significant step forward in teaching machines to understand not just what things are, but how to use them to get a job done.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →