← Latest papers
💻 computer science

EgoIntent: An Egocentric Step-level Benchmark for Understanding What, Why, and Next

This paper introduces EgoIntent, a novel step-level benchmark for egocentric videos that evaluates multimodal large language models on understanding local actions, global intents, and next-step planning, revealing significant challenges in current models' ability to perform fine-grained, anticipatory intent reasoning.

Original authors: Ye Pan, Chi Kit Wong, Yuanhuiyi Lyu, Hanqian Li, Jiahao Huo, Jiacheng Chen, Lutao Jiang, Xu Zheng, Xuming Hu

Published 2026-03-13
📖 5 min read🧠 Deep dive

Original authors: Ye Pan, Chi Kit Wong, Yuanhuiyi Lyu, Hanqian Li, Jiahao Huo, Jiacheng Chen, Lutao Jiang, Xu Zheng, Xuming Hu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are watching a friend fix a bicycle. You see them pick up a wrench, tighten a bolt, and then pause.

  • A basic AI sees: "Person is holding a wrench."
  • A slightly smarter AI sees: "Person is fixing a bike."
  • The "Proactive" AI we want sees: "They are tightening the bolt because the wheel was loose (the 'Why'), and they are about to put the tire back on (the 'What's Next')."

The paper "EgoIntent" introduces a new test to see if current AI models can do that third, most difficult thing: understanding human intent in real-time, step-by-step, without seeing the future.

Here is a breakdown of the paper using simple analogies:

1. The Problem: The "Crystal Ball" Gap

Current AI models are great at watching videos and describing what they see. But they are terrible at being proactive assistants.

If you ask an AI, "What is this person doing?" it can tell you. But if you ask, "Why are they doing this right now, and what will they do next?" it often guesses wrong.

Most existing tests for AI are like watching a whole movie and asking questions at the end. They don't test if the AI can figure out the plot while the movie is still playing. They also often let the AI peek at the ending, which is like giving a student the answer key before the test.

2. The Solution: The "EgoIntent" Benchmark

The researchers created a new test called EgoIntent. Think of it as a driving test for AI, but instead of a car, the AI is a robot watching a human from a first-person perspective (like wearing a GoPro on their head).

The test has three specific challenges, like a three-part puzzle:

  • The "What" (Local Intent): What is the person trying to achieve right this second? (e.g., "They are picking up the drill to make a hole.")
  • The "Why" (Global Intent): What is the bigger goal? (e.g., "They are building a shelf.")
  • The "Next" (Next-Step Plan): What will they do immediately after this step? (e.g., "They will put the screw in the hole.")

3. The Secret Sauce: The "Blindfold" Trick

This is the most important part of the paper. To make the test fair and hard, the researchers use a Temporal Truncation strategy.

Imagine you are watching a magician pull a rabbit out of a hat.

  • Old Tests: Let the AI watch until the rabbit is fully out, then ask, "What was the trick?"
  • EgoIntent: The video cuts off the exact second before the rabbit appears. The AI sees the magician's hand reaching into the hat, but it cannot see the rabbit.

The AI has to guess:

  1. What is the magician doing right now?
  2. Why are they doing it?
  3. What is about to happen?

This prevents the AI from "cheating" by looking at the result. It forces the AI to use anticipation and logic, just like a human does.

4. The Results: The AI is Still a Rookie

The researchers tested 15 different advanced AI models (including the smartest ones from big tech companies) on this benchmark.

The verdict? They all failed miserably.

  • The best AI only got a score of 33 out of 100.
  • Even the smartest models struggled to figure out the "Why" and the "Next" steps.

It's like giving a group of PhD students a simple math problem, and they all get a D. This shows that while AI is getting good at seeing and describing, it is still very bad at understanding human intention and predicting the future.

5. Why Does This Matter?

Why do we care if an AI can't guess what a person will do next?

Because we want AI Assistants and Robots that help us before we ask.

  • Current AI: You say, "I'm thirsty." -> AI gets you water.
  • Future AI (The Goal): You pick up a glass and look at the fridge. -> AI anticipates you are thirsty and opens the fridge for you.

To build robots that can cook with you, help you fix your car, or guide you through surgery, they need to understand the "What, Why, and Next" of every single step. EgoIntent proves that we are still a long way from having robots that can truly "read the room" and think ahead.

Summary Analogy

If current AI is like a tourist who takes photos of a city and writes a postcard about what they saw, EgoIntent is testing if the AI can be a local guide who knows where you are going, why you are going there, and can point out the next turn before you even ask.

Right now, our AI guides are still tourists who are just starting to learn the map.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →