← Latest papers
💻 computer science

LOME: Learning Human-Object Manipulation with Action-Conditioned Egocentric World Model

LOME is an action-conditioned egocentric world model that generates realistic, physically consistent human-object manipulation videos by jointly estimating spatial actions and environmental contexts, thereby outperforming existing methods in motion control and generalization for applications like AR/VR and robotic training.

Original authors: Quankai Gao, Jiawei Yang, Qiangeng Xu, Le Chen, Yue Wang

Published 2026-03-31
📖 5 min read🧠 Deep dive

Original authors: Quankai Gao, Jiawei Yang, Qiangeng Xu, Le Chen, Yue Wang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot how to pour a glass of water, pick up a coffee mug, or stack cups.

In the past, teaching robots (or computers) to do this was like trying to build a car engine from scratch every time you wanted to drive. You had to manually program the physics of gravity, the friction of the table, and the exact shape of every cup. It was slow, rigid, and if you gave the robot a different cup, it would get confused.

LOME (Learning Human-Object Manipulation with Action-Conditioned Egocentric World Model) is a new "magic recipe" that changes the game. Instead of building an engine from scratch, LOME is like a super-smart movie director who has watched millions of hours of human videos and learned how the world works.

Here is how LOME works, broken down into simple concepts:

1. The "First-Person" Perspective (Egocentric)

Most robots see the world like a security camera in the corner of the room. LOME sees the world like you do: through your own eyes. It's "egocentric," meaning it understands that when you reach out with your right hand, the cup moves because you moved it. This perspective is crucial for understanding how we interact with objects in our daily lives.

2. The Three Ingredients

To create a video of a human doing something (like pouring water), LOME needs three specific things, like a chef needing a recipe, ingredients, and a cooking style:

  • The Scene (The Photo): You give it a starting picture (e.g., a table with a green bottle and a mug).
  • The Script (The Text): You tell it what to do (e.g., "Pour the water from the bottle into the mug").
  • The Choreography (The Action Map): This is the secret sauce. Instead of just saying "pour," LOME looks at a map of exactly where your hands are moving at every single second. It's like giving the director a dance routine to follow.

3. The "Joint Dance" (The Secret Sauce)

Here is where LOME gets really clever.

In older methods, the computer would try to guess the video and the hand movements separately. It was like asking a painter to paint a dance while the dancer is standing still in a different room. The result was often messy: the hand might reach for the cup, but the cup wouldn't move, or the water would just appear out of thin air.

LOME does something different. It treats the hand movement and the video as a single, connected dance.

  • The Analogy: Imagine a puppet master (the hand) and a puppet (the object). Old methods tried to animate the puppet based on a description of the master. LOME learns the joint relationship. It understands that because the hand twists the bottle, the water must flow. It learns the physics of the interaction, not just the look of it.

4. What Makes It Special?

  • It Gets the Physics Right: If you ask LOME to pour water, it doesn't just draw a blue stream. It understands that the water level in the mug should rise, and the level in the bottle should fall. It simulates the "cause and effect" of the real world.
  • It's Flexible: If you show it a video of a human picking up a red apple, and then ask it to pick up a blue ball, it can do it. It hasn't memorized the specific apple; it has learned the concept of "picking up."
  • No 3D Modeling Needed: Usually, to make realistic robot movements, you need complex 3D blueprints of every object. LOME skips this. It learns directly from 2D videos, making it much faster and easier to use in the real world.

The Result

When you run LOME, it generates a video that looks like a real person doing the task.

  • The Hands: They move exactly where you told them to.
  • The Objects: They react realistically (liquid flows, cups stack, things fall if dropped).
  • The Consistency: The video doesn't glitch or flicker; it flows smoothly from start to finish.

Why Does This Matter?

Think of LOME as a universal simulator for reality.

  • For Robots: Instead of building a robot in a lab and spending months teaching it how to open a fridge, we can use LOME to generate thousands of "what-if" scenarios (videos) to train the robot. The robot learns by watching LOME's movies.
  • For Virtual Reality (VR/AR): Imagine putting on VR goggles and asking, "Show me how to fix this leaky faucet." LOME could instantly generate a photorealistic video of a human fixing it, right in front of your eyes, so you can learn by watching.

In short, LOME is a bridge between what we say ("pour the water"), what we do (the hand movements), and what happens (the water flows), creating a realistic movie of human interaction without needing a physical set or a 3D blueprint.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →