LongEgoRefer: A Benchmark for Long-Form Egocentric Video Referring Expression Comprehension
This paper introduces LongEgoRefer, a challenging new benchmark constructed from long-form Ego4D videos that addresses the limitations of existing short-clip datasets by featuring extreme target sparsity and complex interactions, revealing that current state-of-the-art models struggle significantly with long-form egocentric spatio-temporal grounding.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are wearing a camera on your head that records your entire day, from the moment you wake up until you go to sleep. This results in a massive, unedited movie that is about 45 minutes long. Now, imagine someone hands you a specific sentence, like: "Find the moment where you pick up the blue mug with the chip on the handle and pour coffee into it."
Your job is to watch that 45-minute movie and pinpoint exactly when that happens and where in the frame the mug is.
This is the challenge the paper "LongEgoRefer" introduces. Here is a breakdown of what they did, using simple analogies:
1. The Problem: The "Needle in a Haystack"
Previous video games for AI were like looking for a needle in a small pile of hay. They used short video clips (a few seconds long) where the object was always visible, like a bright red ball bouncing in a room.
But real life isn't like that. In real life, the "needle" (the object you are looking for) might only appear for a few seconds in a 45-minute movie. It might be hidden behind your hand, or you might only touch it once before putting it away. The paper argues that current AI models are terrible at this because they are used to short, easy clips. They get lost in the long movie and forget what they are looking for.
2. The Solution: A New "Exam" for AI
The authors created a new benchmark called LongEgoRefer. Think of this as a new, much harder final exam for AI models.
- The Source: They took thousands of hours of real-life "first-person" videos (from a dataset called Ego4D) where people are doing daily tasks.
- The Questions: They wrote 1,498 specific questions (called "referring expressions"). These aren't simple like "Where is the cup?" They are complex stories like: "The clear glass jug, which was empty on the scale, was then filled with dark grains from a silver bag."
- The Difficulty:
- Time: The videos are huge (average 45 minutes).
- Sparsity: The object they are asking about might only show up for 1% of the video. It's like asking someone to find a specific 10-second conversation in a 45-minute radio broadcast.
- Complexity: The descriptions require understanding how objects move and interact with hands, not just what they look like.
3. The Test: How Did the AI Do?
The authors tested the smartest AI models available (including giants like GPT-5 and Gemini) on this new exam.
- The Result: The AI models struggled significantly. Even the most advanced models got very low scores.
- The Analogy: It's like giving a student a math test where they have to remember a specific number from a book they read 40 minutes ago, while the book is constantly flipping pages. Most models got lost in the middle of the book.
- The Winner: The model named GPT-5 performed the best, but even it only got about 16% of the "spatio-temporal" (time and place) details correct. This proves the task is incredibly hard, even for the best technology we have right now.
4. Why Did They Fail?
The paper found two main reasons the AI failed:
- Memory Loss: The models couldn't keep track of the whole 45-minute story. They forgot what the object looked like by the time they reached the end of the video.
- The "When" vs. "Where" Trap: The AI has to figure out when the event happens first, and then where the object is. If the AI guesses the wrong time (e.g., "It happened at minute 10" when it actually happened at minute 30), it fails completely, even if it knows what the object looks like. It's like trying to find a specific page in a book but starting your search on the wrong chapter.
5. The Takeaway
The paper concludes that we need a new generation of AI that can truly "watch" long videos without getting confused or forgetting details. They have released this new "exam" (the benchmark) and the code so other researchers can try to build better models to solve this "needle in a haystack" problem.
In short: They built a super-hard test using real-life, long videos to show that today's AI is still bad at finding specific, rare moments in long streams of video, and they are challenging the world to fix it.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.