← Latest papers
💻 computer science

Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding

The paper introduces Minerva-Ego, a benchmark for complex egocentric video reasoning featuring spatiotemporally dense human-annotated traces, which reveals that state-of-the-art models significantly lag behind humans but can be substantially improved by providing "where" and "when" visual hints.

Original authors: Arsha Nagrani, Jasper Uijilings, Shyamal Buch, Tobias Weyand, Sudheendra Vijayanarasimhan, Bo Hu, Ramin Mehran, David A Ross, Cordelia Schmid

Published 2026-05-18
📖 4 min read☕ Coffee break read

Original authors: Arsha Nagrani, Jasper Uijilings, Shyamal Buch, Tobias Weyand, Sudheendra Vijayanarasimhan, Bo Hu, Ramin Mehran, David A Ross, Cordelia Schmid

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a very smart, but slightly distracted, robot how to cook a meal by watching your hands. You show the robot a video of you cooking, then ask a tricky question like, "After I grabbed the oil and opened the drawer, what was the third thing I touched?"

This paper, Minerva-Ego, is about a new "test" the researchers built to see how good these robots (AI models) are at answering such questions. But here's the twist: they didn't just grade the robots on whether they got the final answer right. They also gave them a "cheat sheet" showing exactly where and when to look in the video to find the answer.

Here is the breakdown of what they did and found, using simple analogies:

1. The Problem: The Robot is a "Distracted Student"

Current AI models are like students who can read a book but get lost if you ask them to find a specific detail in a 30-minute movie.

  • The Old Way: Previous tests only asked, "What is the answer?" If the robot said "Knife," it got a point. If it said "Spoon," it got zero. The test didn't care why the robot got it wrong.
  • The New Way (Minerva-Ego): The researchers created a test where every question comes with a detailed map. This map tells the robot: "At 12:56, look at the oil bottle. At 12:58, look at the drawer. At 13:04, look at the spoon."
  • The Result: When they tested the smartest robots available (like Gemini and GPT), they found the robots were still struggling. They got the final answer right only about 30-40% of the time, while humans got it right 92% of the time.

2. The Diagnosis: "I Can't Find the Needle in the Haystack"

The researchers looked closely at why the robots failed. They found two main issues:

  • Perceptual Blindness: The robot couldn't "see" the right object. It might think a pan is a pot, or it misses an object entirely because the camera angle (first-person view) is tricky.
  • Time Confusion: The robot gets the timeline mixed up. It might remember the spoon but forget when it happened, mixing up the order of events.

3. The Solution: The "Flashlight" and the "Highlighter"

To fix this, the researchers tried a new trick. Instead of just showing the whole video, they gave the robot Spatio-Temporal Hints.

  • Spatial Hint (The Flashlight): They drew a red circle or a box around the specific object the robot needed to see (like the spoon or the oil bottle). It's like shining a flashlight on the exact spot the robot needs to look at.
  • Temporal Hint (The Highlighter): They didn't show the robot the whole 10-minute video. Instead, they only showed the specific 5 seconds where the spoon appeared. It's like highlighting the exact paragraph in a book instead of making the student read the whole chapter.

4. The Results: A Big Jump in Performance

When they gave the robots these hints:

  • The "Oracle" Test (Perfect Hints): When they used perfect, human-drawn maps to tell the robot exactly where to look, the robot's score jumped by about 5.6%. It proved that if the robot could just find the right object at the right time, it could reason much better.
  • The "Real World" Test (AI Hints): They also tried using a standard AI tool to draw the boxes automatically (without human help). It still helped, though not quite as much as the perfect human maps.

5. The Big Takeaway

The paper concludes that the main reason robots are bad at understanding these videos isn't that they can't "think" or "reason" logically. They actually have decent logic. The problem is perception. They are failing to "look" at the right thing at the right time.

In a nutshell:
Imagine you are playing a game of "Where's Waldo?" but the book is 100 pages long and Waldo moves around.

  • Old AI: Tries to scan the whole book, gets tired, and guesses "Waldo is on page 50." (Wrong).
  • Minerva-Ego: Says, "Hey, look at page 52, row 3, column 2. There he is."
  • Result: The AI suddenly gets the answer right.

The paper proves that for robots to get better at understanding our daily lives (embodied agents), we don't just need smarter brains; we need better ways to help them focus their eyes on the right things at the right moments.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →