← Latest papers
💬 NLP

StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos

This paper introduces StreamGaze, the first benchmark designed to evaluate how Multimodal Large Language Models leverage human gaze signals for temporal reasoning and proactive intention prediction in streaming videos, revealing significant performance gaps between current models and human capabilities.

Original authors: Daeun Lee, Subhojyoti Mukherjee, Branislav Kveton, Ryan A. Rossi, Viet Dac Lai, Seunghyun Yoon, Trung Bui, Franck Dernoncourt, Mohit Bansal

Published 2026-03-30
📖 5 min read🧠 Deep dive

Original authors: Daeun Lee, Subhojyoti Mukherjee, Branislav Kveton, Ryan A. Rossi, Viet Dac Lai, Seunghyun Yoon, Trung Bui, Franck Dernoncourt, Mohit Bansal

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are wearing a pair of smart AR glasses that are trying to help you cook dinner. You want these glasses to be the perfect sous-chef: they should know what you are looking at, remember what you saw earlier, and even predict what you are about to do next.

The paper "STREAMGAZE" introduces a new way to test if AI computers are smart enough to be these sous-chefs. Here is the breakdown in simple terms:

1. The Problem: The AI is "Blind" to Your Eyes

Current AI video models are like a security camera that records everything but doesn't know what you care about.

  • The Scenario: You are cooking. You look at a pot, then a knife, then a tomato.
  • The AI's Mistake: Without "eye-tracking," the AI just sees a blur of kitchen stuff. It doesn't know you are focused on the pot and ignoring the refrigerator.
  • The Gap: Previous tests checked if AI could understand video, but they didn't check if the AI could understand your attention.

2. The Solution: STREAMGAZE (The "Eye-Test" for AI)

The researchers built a massive new test called STREAMGAZE. Think of it as a driver's license exam for AI, but instead of checking if the AI can drive a car, they check if it can "drive" by following your eyes in real-time.

They created over 8,500 questions based on real videos of people cooking, working in labs, and assembling things. The test has three levels of difficulty, just like a video game:

  • Level 1: The Memory Game (Past Tasks)

    • The Question: "You looked at the knife 10 seconds ago. What was sitting on the counter next to it that you didn't look at?"
    • The Challenge: The AI must remember the background details of what you saw, even if you weren't staring directly at them.
  • Level 2: The "What's Happening Now?" Game (Present Tasks)

    • The Question: "What is the user looking at right this second? Is it the red pepper or the green one?"
    • The Challenge: The AI must pinpoint exactly where your eyes are focused in a moving, shaky video.
  • Level 3: The Crystal Ball (Proactive Tasks)

    • The Question: "The user just looked at the water bottle and the stove. Should I alert them that the water is boiling?" or "What will they do next?"
    • The Challenge: The AI has to guess your intent before you even move. It's like a GPS that says, "You're looking at the exit, so you probably want to leave," before you even say it.

3. How They Built It: The "Gaze-Mapper"

To make this test, the researchers had to teach a computer to translate "eye movements" into "video understanding."

  • The Process: They took videos of people wearing eye-trackers. They found the moments where the person's eyes stopped moving (called fixations)—like when you stare at a tomato to chop it.
  • The Map: They drew a circle around where the eyes were looking (the FOV or Field of View).
  • The Translation: They used a super-smart AI to label everything inside that circle (the tomato) and everything outside it (the fridge in the background).
  • The Result: A "Scanpath"—a story of where the eyes went, step-by-step, like a treasure map of attention.

4. The Shocking Results: AI is Still a Rookie

When they tested the world's best AI models (like GPT-4o and others) on STREAMGAZE, the results were humbling.

  • Human Score: Humans got about 83% correct. We are naturally good at knowing what we are looking at and what we might do next.
  • AI Score: The best AI models only got about 45-50% correct.
  • The Verdict: The AI is terrible at "reading the room." It gets confused by the flow of time. It often forgets what it saw 5 seconds ago or fails to predict that if you look at a knife, you are probably about to cut something.

5. Why This Matters

This isn't just about cooking videos. This is about the future of Augmented Reality (AR) glasses and robots.

  • If you wear AR glasses, you want them to highlight the tool you are looking at and hide the clutter you aren't.
  • If you have a robot assistant, you want it to hand you the salt before you ask for it, because it saw you looking at the soup.

STREAMGAZE proves that while AI is getting better at "seeing" videos, it is still very bad at "understanding" human attention. It's like having a very smart librarian who can read every book in the library but doesn't know which book you are trying to find.

In short: The paper says, "We built a new, harder test for AI that includes eye-tracking. The AI failed the test, showing us exactly where we need to improve to build truly helpful, human-like assistants."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →