← Latest papers
💻 computer science

Do MLLMs Understand Pointing? Benchmarking and Enhancing Referential Reasoning in Egocentric Vision

This paper introduces EgoPoint-Bench, a comprehensive benchmark comprising over 11,000 samples to evaluate and address "Referential Hallucination" in Multimodal Large Language Models, demonstrating that fine-tuning on their synthetic data significantly enhances precise spatial grounding for egocentric pointing gestures.

Original authors: Chentao Li, Zirui Gao, Mingze Gao, Yinglian Ren, Jianjiang Feng, Jie Zhou

Published 2026-04-24
📖 4 min read☕ Coffee break read

Original authors: Chentao Li, Zirui Gao, Mingze Gao, Yinglian Ren, Jianjiang Feng, Jie Zhou

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are wearing a pair of smart glasses that act as your personal AI assistant. You're walking through a cluttered room, and you point your finger at a shelf to ask, "What is this used for?"

Ideally, the AI should look at exactly where your finger is pointing and tell you, "That's a bookshelf, used for holding books."

But in reality, today's smart AI glasses are like a confused tourist. Instead of looking at where your finger is pointing, they just look at whatever is closest to your hand or whatever object is the brightest and most colorful in the room. If you point at a bookshelf, but there's a shiny red apple nearby, the AI might say, "That's an apple!" because it got distracted by the apple's color rather than following your finger.

The researchers in this paper call this "Referential Hallucination." It's like the AI is having a daydream about what it thinks you mean, rather than understanding what you are actually pointing at.

The Problem: The "Pointing" Gap

The paper argues that while AI has gotten very good at describing pictures (like saying, "There is a cat on a mat"), it is terrible at understanding pointing gestures from a first-person perspective (like a human wearing glasses).

Current AI models rely on "lazy shortcuts." They assume that if you point near an object, you probably mean that object. But in the real world, your finger might be pointing past a chair to a lamp behind it. The AI fails to trace the invisible "laser beam" of your finger to the actual target.

The Solution: Building a "Pointing Gym"

To fix this, the team created a massive training ground called EgoPoint-Bench. Think of this as a gym for AI eyes.

  1. The Simulation (The Virtual Gym):
    They built a super-realistic 3D video game world. Inside this world, they programmed thousands of virtual hands to point at thousands of different objects (chairs, toasters, fire extinguishers).

    • Why a video game? Because in the real world, it's hard to get perfect data. In the game, they know exactly where the finger is pointing. They can draw an invisible line from the fingertip to the object and say, "This is the truth."
    • They generated over 10,000 of these perfect examples, covering everything from pointing at a specific book to pointing at a vague spot in a room.
  2. The Real World (The Field Test):
    They also recorded real people wearing smart glasses in real places (stores, homes, zoos) to make sure the AI could handle messy, real-life situations.

  3. The Training (The Workout):
    They took existing smart AI models and "exercised" them using this new dataset. They taught the AI: "No, don't look at the apple next to the finger! Look at the book the finger is actually touching!"

The Results: From Confused to Sharp

When they tested the AI before this training, it was like a student who failed the test, getting about 60% of the answers right. It was guessing based on what looked "important" rather than what was being pointed at.

After the "workout" (fine-tuning on their data):

  • The AI's accuracy jumped significantly (up to 75-80%).
  • The Magic Trick: Even though the AI was only trained on the video game data, it got much better at understanding real-world pointing. It's like practicing your tennis swings in a simulator and then suddenly playing better on a real court. The AI learned the logic of pointing, not just the look of the objects.

Why This Matters

This research is a huge step toward making Augmented Reality (AR) and smart glasses actually useful.

Right now, if you ask your smart glasses a question while pointing, they might misunderstand you. This paper provides the roadmap to teach AI how to truly "see" what you are pointing at, making future interactions feel natural and intuitive, just like talking to a human friend who understands your gestures.

In short: The authors built a giant, perfect practice field to teach AI how to follow a finger, fixing a major blind spot in how computers understand human gestures.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →