← Latest papers
💻 computer science

MomaGraph: State-Aware Unified Scene Graphs with Vision-Language Model for Embodied Task Planning

The paper introduces MomaGraph, a unified, state-aware scene graph representation and a corresponding large-scale dataset and evaluation suite, alongside MomaGraph-R1, a vision-language model that leverages these resources to achieve state-of-the-art performance in task-oriented scene graph prediction and embodied task planning.

Original authors: Yuanchen Ju, Yongyuan Liang, Yen-Jen Wang, Nandiraju Gireesh, Yuanliang Ju, Seungjae Lee, Qiao Gu, Elvis Hsieh, Furong Huang, Koushil Sreenath

Published 2026-02-10
📖 4 min read☕ Coffee break read

Original authors: Yuanchen Ju, Yongyuan Liang, Yen-Jen Wang, Nandiraju Gireesh, Yuanliang Ju, Seungjae Lee, Qiao Gu, Elvis Hsieh, Furong Huang, Koushil Sreenath

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a guest in a massive, high-tech mansion. You want to make a cup of tea, but you’ve never been there before. You see a kitchen, but you don't know which knob turns on the stove, which button starts the kettle, or if the "handle" you see is for a drawer or a cupboard.

To a standard robot, this mansion is just a giant, confusing pile of shapes and colors. To solve this, researchers have created MomaGraph, which essentially gives a robot a "mental map" that is much smarter than a simple GPS.

Here is the breakdown of how it works using everyday analogies:

1. The Problem: The "Blind Tourist" Robot

Most current robots are like tourists looking through a tiny straw. They can see an object (like a stove), but they don't understand its "personality."

  • Spatial knowledge is knowing where the stove is (it's in the corner).
  • Functional knowledge is knowing how it works (this knob controls the heat).

Current robots usually only have one or the other. They might know where the stove is, but they try to "open" it like a door because they don't understand it has knobs. They are like someone who knows where the library is but doesn't realize you have to open a book to read it.

2. The Solution: MomaGraph (The "Smart Blueprint")

MomaGraph is like giving the robot a "Smart Blueprint" of the room. Instead of just a map of walls, this blueprint has "sticky notes" on everything.

  • It connects the dots: It doesn't just say "there is a knob" and "there is a stove." It draws a line between them that says: "This specific knob is the boss of this specific burner."
  • It looks at the details: It doesn't just see a "fridge"; it sees the "handle" as a specific part you need to grab.
  • It’s Task-Focused: If you tell the robot "Make tea," it doesn't waste brainpower thinking about the sofa. It "zooms in" on the kettle, the water, and the stove, just like a human does.

3. The Brain: MomaGraph-R1 (The "Experienced Chef")

To make this blueprint, they built a brain called MomaGraph-R1. They trained it using something called Reinforcement Learning.

Think of this like training a puppy or a chef. Instead of just showing the robot a thousand pictures of kitchens (which is what most AI does), they gave it a "reward" every time it drew an accurate blueprint.

  • If the robot drew a map that correctly linked the remote to the TV, it got a "digital treat."
  • If it missed a step (like trying to boil water before plugging in the kettle), it didn't get a treat.
    Over time, the robot stopped just "guessing" what things looked like and started "reasoning" about how they work.

4. The "Aha!" Moment: State-Awareness (The "Detective")

One of the coolest parts is that the robot is a detective.
Imagine a kitchen with four identical stove knobs. The robot doesn't know which one is the right one. But in MomaGraph, the robot can try one, see that the flame turns on, and say, "Aha! I've solved the mystery. This knob is the controller!" It updates its mental map in real-time as it interacts with the world. It learns from its own actions.

Summary: Why does this matter?

In short, MomaGraph moves robots from being "clumsy movers" (who can move from point A to point B) to "capable helpers" (who understand that to turn on the light, they must find the switch, not the lamp).

It gives them the ability to look at a messy, complicated room and say: "I see the goal, I see the tools, and I know exactly how they connect to get the job done."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →