← Latest papers
💻 computer science

Do You See What I Am Pointing At? Gesture-Based Egocentric Video Question Answering

This paper introduces EgoPointVQA, a new dataset and benchmark for gesture-grounded egocentric question answering, along with the Hand Intent Tokens (HINT) method that leverages 3D hand keypoints to significantly improve multimodal large language models' ability to infer pointing intent from egocentric videos.

Original authors: Yura Choi, Roy Miles, Rolandos Alexandros Potamias, Ismail Elezi, Jiankang Deng, Stefanos Zafeiriou

Published 2026-03-17
📖 3 min read☕ Coffee break read

Original authors: Yura Choi, Roy Miles, Rolandos Alexandros Potamias, Ismail Elezi, Jiankang Deng, Stefanos Zafeiriou

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are wearing a pair of smart glasses that can see exactly what you see. You want to ask your AI assistant a question about the world around you.

The Problem: The "Pointing" Gap
Right now, if you point at a specific pot on your stove and ask, "Should I use this one?", most AI assistants get confused. They might look at the whole kitchen, see two pots, and guess randomly. They don't understand that your finger is the "cursor" telling them exactly which pot you mean.

It's like trying to play a game of "I Spy" with someone who is blindfolded; you say, "I spy something red," but they can't see your finger pointing at the specific red apple on the table, so they guess the red car in the driveway instead.

The Solution: EGOPOINTVQA
The researchers behind this paper built a new training ground called EGOPOINTVQA. Think of this as a massive "driving school" for AI, but instead of learning to drive cars, the AI is learning to understand pointing.

They created thousands of videos (some made by computers, some filmed by real people wearing smart glasses) where someone points at objects and asks questions like:

  • "What is this?"
  • "Which of these is closer?"
  • "How many of these are there?"

The goal was to teach the AI that when a human points, the word "this" or "that" isn't just a random word; it's a command that says, "Look right here, at the thing my finger is touching."

The Secret Sauce: HINT (Hand Intent Tokens)
Even with all these videos, standard AI models still struggled. They could see the video and read the question, but they couldn't "feel" the gesture.

To fix this, the researchers invented a clever trick called HINT (Hand Intent Tokens).

Imagine you are trying to explain a complex dance move to a friend over the phone. You could just describe the steps (the text), or you could send them a video of the dance (the visual). But what if you sent them a special map that highlights exactly where your hands are moving at every second?

That's what HINT does.

  1. The Map: The system uses a special tool to track the 3D position of the user's hand and fingers in every single frame of the video.
  2. The Translation: It turns those hand positions into a secret code (tokens) that the AI can read.
  3. The Mix: It mixes this "hand code" right into the AI's brain alongside the video and the question.

Now, when the AI reads the question "Is this pot black?", it doesn't just guess. It looks at the "hand map," sees the finger pointing at the black pot, and confidently answers, "Yes."

Why This Matters
This research is a big step toward making AI assistants that feel truly natural. In the future, you won't have to say, "Hey AI, look at the black pot on the left." You can just point at it and ask, "Is this one ready?" and the AI will know exactly what you mean.

In a Nutshell:

  • The Issue: AI is bad at understanding when humans point at things.
  • The Dataset: They built a giant library of pointing videos to train AI.
  • The Fix: They gave the AI a "superpower" (HINT) that translates hand movements into a language the AI understands, allowing it to finally "see" what you are pointing at.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →