← Latest papers
💻 computer science

FAWAM: Force-Aware World Action Models for Closed-Loop Contact-Rich Manipulation

FAWAM is a force-aware world action model that integrates 6-axis force/torque signals across perception, prediction, and closed-loop execution to explicitly model contact evolution and refine actions online, achieving significantly higher success rates in contact-rich robotic manipulation compared to vision-only and existing force-aware baselines.

Original authors: Haotian He, Zeyu Yan, Qipeng Liu, Ning Guo, Wenzhao Lian

Published 2026-06-09
📖 4 min read☕ Coffee break read

Original authors: Haotian He, Zeyu Yan, Qipeng Liu, Ning Guo, Wenzhao Lian

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a robot trying to wipe a whiteboard, peel a cucumber, or push a box. To a human, these tasks feel natural because we have a "sixth sense" for touch. We know instantly if we are pressing too hard, if something is slipping, or if we've hit a bump. We don't just look at the object; we feel the interaction and adjust our grip or pressure in real-time.

Current robots, however, are often like people trying to do these tasks while wearing thick, numb gloves and blindfolds. They rely mostly on cameras (vision). If the camera sees a whiteboard, the robot tries to wipe it. But if the robot presses too hard and the marker smears, or if it presses too lightly and nothing happens, the robot often doesn't know why it failed until it's too late. It's like trying to drive a car by only looking at a map, ignoring the feeling of the steering wheel or the sound of the engine.

FAWAM is a new "brain" for robots that gives them back that sense of touch, but with a superpower: it can predict the future.

Here is how FAWAM works, broken down into three simple steps using a cooking analogy:

1. The "Chef's Intuition" (Perception)

Most robots just look at the ingredients (vision). FAWAM is like a chef who not only sees the pot but also feels the heat rising and smells the steam.

  • What it does: It takes the history of forces (how hard the robot pushed or pulled in the last few seconds) and uses that to understand the current situation.
  • The Analogy: Instead of just seeing a pot boiling, the robot "feels" the vibration and heat to know exactly how the soup is reacting to the spoon. This helps it understand the "contact state"—is the object sliding? Is it stuck?

2. The "Crystal Ball" (Prediction)

This is the paper's biggest innovation. Most robots react to what is happening now. FAWAM asks, "If I do this action, what will happen to the force in the next second?"

  • What it does: It doesn't just predict where the robot's hand will go; it predicts the force the hand will feel. It simulates the future interaction.
  • The Analogy: Imagine a chef tasting a spoonful of soup and immediately predicting, "If I add more salt, it will be too salty in 10 seconds." FAWAM does this with physics. It says, "If I push this box forward, I predict I will feel a sudden resistance in 0.5 seconds." By predicting the future "ouch" or "slip," the robot can plan better.

3. The "Instant Reflex" (Correction)

Even with a crystal ball, things can go wrong. Maybe the table is slightly uneven, or the object is heavier than expected.

  • What it does: While the robot is moving, it constantly compares its prediction (what it thought would happen) with reality (what the sensors actually feel). If there is a mismatch, it instantly tweaks its movement to fix it.
  • The Analogy: This is like a tightrope walker. They have a plan to walk across, but if they feel a sudden gust of wind (a deviation from their prediction), they instantly shift their weight to stay balanced. FAWAM does this 10 times a second. If the robot predicted a smooth slide but feels a sudden jerk, it instantly adjusts its grip to prevent a crash.

The Results: Why It Matters

The researchers tested this on real robots doing messy, touch-heavy tasks like peeling a cucumber or wiping a vase.

  • Vision-only robots (the "blindfolded" ones) succeeded less than half the time.
  • Robots that just "felt" the force (without the prediction and correction) did better, but still struggled.
  • FAWAM succeeded 85% of the time.

The paper claims that by combining feeling the past, predicting the future force, and correcting in the moment, the robot becomes much more robust. It doesn't just react to a problem; it anticipates it and fixes it before the task fails.

In short: FAWAM turns a robot from a clumsy, sight-only machine into a skilled artisan that can "feel" its way through complex tasks, predicting bumps before they happen and adjusting its grip instantly to keep the job done.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →