← Latest papers
💻 computer science

Prime and Reach: Synthesising Body Motion for Gaze-Primed Object Reach

This paper introduces a novel dataset of 23.7K gaze-primed motion sequences and a text-conditioned diffusion model that generates realistic full-body motions, successfully capturing the natural human behavior of spotting an object before reaching it.

Original authors: Masashi Hatano, Saptarshi Sinha, Jacob Chalk, Wei-Hong Li, Hideo Saito, Dima Damen

Published 2026-03-26
📖 4 min read☕ Coffee break read

Original authors: Masashi Hatano, Saptarshi Sinha, Jacob Chalk, Wei-Hong Li, Hideo Saito, Dima Damen

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are walking into a kitchen to grab a jar of peanut butter from a high shelf. Before you even lift your hand, your eyes have already locked onto that jar. You might even tilt your head slightly or shift your weight to get a better angle. This split-second moment of "looking before reaching" is called priming. It's your brain's way of saying, "I see the target, and I'm getting ready to grab it."

For a long time, computers trying to animate human movement (like in video games or robots) were terrible at this. They could make a character walk or run, but when it came to grabbing something, the character would often just teleport their hand to the object or reach blindly without looking first. It looked robotic and unnatural.

This paper, "Prime and Reach," is like a new recipe for teaching computers how to be more human. Here is the breakdown:

1. The Problem: The "Blind Grab"

Imagine a robot trying to pick up a cup. If you just tell it, "Go get the cup," the robot might spin around and slam its hand into the cup without ever looking at it. It lacks the anticipation that humans have. Humans don't just move; we plan our movement by looking at where we are going first.

2. The Solution: A Massive Library of "Looking"

The researchers realized they needed to teach the computer what "looking before grabbing" actually looks like. So, they went on a scavenger hunt through five different video datasets (collections of real people doing daily tasks).

  • The Hunt: They looked for moments where a person's eyes landed on an object before their hand touched it.
  • The Result: They curated a massive library of 23,700 of these specific "look-and-grab" sequences. Think of this as a giant cookbook of "how to look before you reach."

3. The Teacher: The "Diffusion" Model

To teach the computer, they used a type of AI called a Diffusion Model.

  • The Analogy: Imagine a sculpture made of clay that is currently just a messy pile of noise (like static on an old TV). The AI acts like a sculptor who slowly chips away the noise, step-by-step, until a clear, smooth statue emerges.
  • The Twist: In this case, the AI isn't just chipping away noise; it's being guided by a "goal." The researchers told the AI: "Start with this messy noise, but make sure the final statue looks like a person who looked at a cup before grabbing it."

4. The Two Ways to Give Instructions

The researchers tested two ways to tell the AI what to do:

  1. The "Pose" Instruction: "Stand here, and end up in this exact pose holding the cup." (Like giving a dancer a specific final pose).
  2. The "Location" Instruction: "The cup is over there. Go get it." (This is harder because the AI has to figure out how to stand and move to get there).

5. The Magic Metric: "Prime Success"

How do you know if the robot is doing it right? They invented a new test called Prime Success.

  • The Test: Did the character's eyes (or head direction) actually lock onto the object before the hand touched it?
  • The Result: Their new AI scored incredibly high on this test. It learned to "look" at the target, pause for a split second to orient itself, and then reach. It stopped doing the "blind grab."

6. Why This Matters

This isn't just about making better video games.

  • For Robots: If a robot is helping you in a kitchen or a hospital, it needs to look at what it's touching to be safe and efficient.
  • For Virtual Humans: If you are in a virtual reality meeting, you want the avatar to look at you when you speak, not stare at the ceiling while reaching for a coffee cup.

The Bottom Line

The researchers took a complex human behavior (the subtle art of looking before reaching), gathered thousands of real-life examples, and taught an AI to mimic it. The result is a digital human that doesn't just move; it anticipates. It looks at the destination before it moves, making the movement feel natural, fluid, and truly human.

In short: They taught the computer that you have to look at the cookie before you grab it.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →