← Latest papers
💻 computer science

Point & Grasp: Flexible Selection of Out-of-Reach Objects Through Probabilistic Cue Integration

This paper introduces "Point & Grasp," a mixed reality interaction technique that uses a probabilistic framework to flexibly integrate multiple user cues—specifically pointing direction and grasp gestures—to more accurately and robustly select out-of-reach objects.

Original authors: Xuejing Luo, Hee-Seung Moon, Christian Holz, Antti Oulasvirta

Published 2026-04-27
📖 4 min read☕ Coffee break read

Original authors: Xuejing Luo, Hee-Seung Moon, Christian Holz, Antti Oulasvirta

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are playing a high-tech video game or working in a virtual office, and you see a tool you need—like a screwdriver—floating high up on a shelf or far across the room. You can’t physically reach it, so you have to "point and click" to grab it.

In current Virtual Reality (VR), this is actually surprisingly annoying. It’s like trying to pick a specific grape out of a crowded bowl using only a long, shaky laser pointer. If the grapes are too close together, your laser hits the wrong one. If you try to use a "grabbing" motion, but all the fruits look similar, the computer gets confused.

This paper introduces a new way to solve this called Point&Grasp.

The Problem: The "Shaky Laser" and the "Generic Grab"

The researchers identified two main ways people currently try to select things in VR, and both have "blind spots":

  1. The Directional Cue (The Shaky Laser): You point your hand at the object. This works great if the object is alone, but if there’s a cluster of items, your "laser" is too imprecise. It’s like trying to point at a single hair on a cat's head from across the room.
  2. The Gestural Cue (The Generic Grab): You make a shape with your hand that looks like you're grabbing something. This is great for telling the computer what you want (e.g., "I'm making a shape for a handle"), but if you have three different tools with similar handles, the computer doesn't know which one you mean.

The Solution: The "Smart Detective" (Bayesian Integration)

Instead of forcing the computer to choose between the "laser" or the "grab," the researchers created a "Smart Detective" (technically called a Probabilistic Framework).

Think of it like this: Imagine you are in a dark room and you want to find a specific coffee mug.

  • Cue 1 (The Laser): You shine a flashlight in a general direction. It hits a cluster of three mugs. You know it's one of those three, but you aren't sure which.
  • Cue 2 (The Grab): You hold your hand in a specific way, like you're gripping a thin handle.

The Smart Detective doesn't just look at the flashlight or just look at your hand. It combines the clues. It says: "Okay, the flashlight is pointing at these three mugs, but the hand shape only really fits the handle of the blue one. Therefore, there is a 95% chance they want the blue mug."

By mathematically "fusing" these two clues, the system becomes much more accurate and much faster.

How they built it: The "Training Camp"

To make this detective smart, they couldn't just use old data. Most VR data is about people grabbing things they can actually touch. But when you grab something out of reach, your hand behaves differently—you tend to stretch your fingers more or hold your hand differently because you're "imagining" the touch.

The researchers created a massive new dataset called ORG (Out-of-Reach Grasping). They had people in VR practice "imaginary grabbing" hundreds of times. They recorded every finger position and every object shape. They then used this "training camp" to teach a neural network how to recognize the difference between a "mug grab" and a "hammer grab" in mid-air.

The Result: A Smoother Virtual World

When they tested Point&Grasp against the old methods, the results were clear:

  • It’s faster: You don't spend time "fine-tuning" your pointer.
  • It’s more reliable: Even when objects are crowded together or look similar, the "Smart Detective" usually figures it out.
  • It feels natural: Users felt like they were interacting with the world the way they do in real life—using both direction and shape to communicate intent.

In short: Instead of making you be a perfect marksman with a laser pointer, Point&Grasp lets you be a human, combining your pointing and your gesturing to tell the computer exactly what you want.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →