← Latest papers
🤖 AI

ToG-Bench: Task-Oriented Spatio-Temporal Grounding in Egocentric Videos

This paper introduces ToG-Bench, the first task-oriented spatio-temporal video grounding benchmark for egocentric videos, which addresses the limitations of existing object-centric approaches by featuring task-oriented, explicit-implicit dual, and one-to-many grounding scenarios to better evaluate and advance embodied intelligence.

Original authors: Qi'ao Xu, Tianwen Qian, Yuqian Fu, Kailing Li, Yang Jiao, Jiacheng Zhang, Xiaoling Wang, Liang He

Published 2026-04-07
📖 4 min read☕ Coffee break read

Original authors: Qi'ao Xu, Tianwen Qian, Yuqian Fu, Kailing Li, Yang Jiao, Jiacheng Zhang, Xiaoling Wang, Liang He

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a robot to help you around your house. You don't just want the robot to say, "I see a red cup on the table." You want it to understand your intent. If you say, "I'm thirsty, get me some water," the robot needs to figure out that it must find the sink, turn on the faucet, and grab a glass—even though you never explicitly mentioned the faucet or the glass.

This is exactly what the paper ToG-Bench is about. It's a new "exam" designed to test how well AI robots can understand human tasks from a first-person perspective (like wearing a camera on your head).

Here is a breakdown of the paper using simple analogies:

1. The Problem: The "Descriptive" Robot vs. The "Task" Robot

Previous AI tests were like playing a game of "I Spy."

  • The Old Way: You say, "Find the blue pillow." The robot looks for something blue and fluffy. It's easy because the robot just matches words to pictures.
  • The Real World: You say, "Make coffee." The robot needs to know that "making coffee" involves a machine, beans, a filter, and a mug. It has to figure out what to touch and when to touch it, even if you didn't list every single item.

The authors realized that current AI is great at "I Spy" but terrible at "Make coffee." It doesn't understand the goal, only the description.

2. The Solution: ToG-Bench (The New Exam)

The team created a new benchmark called ToG-Bench (Task-Oriented Grounding). Think of this as a driving test for robots. Instead of just asking them to "spot a stop sign," they ask them to "navigate to the grocery store and buy milk."

The exam has three tricky rules:

  • The "Hidden Clue" Rule (Explicit vs. Implicit): Sometimes you say "Turn on the light" (Explicit: the light switch). Sometimes you say "It's dark in here" (Implicit: the robot must realize it needs to find the light switch, even though you didn't say "switch"). The robot must use common sense to fill in the blanks.
  • The "Teamwork" Rule (One-to-Many): A single instruction often needs multiple tools. If you say "Clean the whiteboard," the robot needs to find the board, the marker, and the eraser. It can't just find one; it needs the whole crew.
  • The "Time Travel" Rule (Spatio-Temporal): The robot doesn't just need to find the object; it needs to know when to interact with it. It has to track the object as you move around the room, not just in a single snapshot.

3. How They Built It (The Assembly Line)

Building this exam was hard because you can't just ask a human to write 2,700 complex instructions for 100 videos. That would take forever.

  • The Pipeline: They used a "smart assistant" (a powerful AI called Gemini) to watch videos and suggest tasks like "Wash your hands."
  • The Safety Check: Then, real humans stepped in as editors. They checked: "Is the faucet actually visible? Is the instruction safe? Did the robot track the water correctly?"
  • The Result: A high-quality dataset of 100 videos with over 2,700 tasks, covering everything from simple actions to complex, multi-step chores.

4. The Results: The "Smart but Clumsy" Robot

The authors tested the world's best AI models (like GPT-5 and Gemini) on this new exam. The results were a mix of good news and bad news:

  • The Good News: The AIs are very smart at understanding the idea. If you say "Make coffee," they can usually tell you, "Okay, I need a coffee machine." They are great at the "What" and the "Why."
  • The Bad News: They are terrible at the "Where" and "When."
    • The Drift: Imagine a robot trying to hold a cup while walking. The AI knows it's holding a cup, but as the camera moves, the AI loses track. It might point to the wall instead of the cup.
    • The "Hidden" Struggle: When the task required the robot to guess an object you didn't mention (like the faucet), the AI's performance dropped significantly.
    • The Complexity Crash: When asked to find three things at once (like a computer, mouse, and keyboard), the AI got confused and failed to find them all correctly.

The Big Takeaway

This paper is a wake-up call. We have built AIs that can read a recipe and understand the story, but we haven't taught them how to actually cook in a real kitchen while moving around.

ToG-Bench is the first step in fixing this. It shows us exactly where the robots are failing so engineers can build the next generation of AI that doesn't just "see" the world, but truly understands how to interact with it to get things done.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →