← Latest papers
💻 computer science

BoxTwin: Learning Elastoplastic Articulated Object Dynamics from Videos

BoxTwin is an interactive digital twin framework that learns the full elastoplastic dynamics of articulated objects from videos, enabling robots to accurately track joint trajectories and predict long-term plastic deformation and damage during physical interactions.

Original authors: Heng Zhang, Gehan Zheng, Kaifeng Zhang, Jay Song, Shivansh Patel, Sonny Hu, Yunzhu Li, Changxi Zheng, Peter Yichen Chen

Published 2026-07-21
📖 4 min read☕ Coffee break read

Original authors: Heng Zhang, Gehan Zheng, Kaifeng Zhang, Jay Song, Shivansh Patel, Sonny Hu, Yunzhu Li, Changxi Zheng, Peter Yichen Chen

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot how to fold a piece of cardboard or close a flimsy metal door. If you ask a standard robot to do this, it might treat the object like a solid brick or a perfect spring. It would try to bend the cardboard, but because real-world materials are messy, the robot would get confused. The cardboard might bend back, stay bent, or even crack, and the robot wouldn't know why. This is the world of "digital twins"—a fancy term for creating a perfect, virtual copy of a real object inside a computer so a robot can practice before touching the real thing. But until now, these virtual copies have been terrible at handling things that are part rigid, part squishy, and part broken. They struggle with objects that bend like rubber, stay bent like wet clay, and get weaker every time you fold them. This paper tackles that specific messiness, aiming to give robots a brain that understands how real-world, bendy, breakable things actually behave.

The researchers behind this work, Heng Zhang and their team, have built a new system called BoxTwin. Think of BoxTwin as a super-smart video detective that watches a robot play with a bendy object and learns exactly how that object behaves. Instead of guessing the rules, BoxTwin watches the video, figures out the hidden physics, and builds a digital twin that can predict the future.

Here is the magic trick: most digital twins assume that if you bend something, it snaps back to its original shape (like a rubber band). BoxTwin knows that's not always true. It understands three tricky things that happen when you mess with objects like cardboard boxes or sheet-metal assemblies:

  1. Non-linear Elasticity: Sometimes, the harder you push, the more it resists in a weird, non-straight way.
  2. Plastic Deformation: Sometimes, if you push too hard, the object stays bent forever. It doesn't snap back.
  3. Damage: Every time you bend it, the material gets a little bit weaker, like a paperclip that eventually snaps after too many folds.

BoxTwin doesn't just guess these rules; it learns them. The system takes a video of a real-world object being folded or manipulated. It then reconstructs the scene and figures out the specific "constitutive model" (a fancy way of saying the rulebook for how that specific material moves) for that object. It tracks how the "rest angle" (where the object wants to sit) changes as it gets bent, and it even counts how much "damage" has accumulated to make the object weaker over time.

To test if this actually works, the team ran two experiments. First, they manually folded a simple object made of two panels and a joint, keeping one side stuck to a table. They used colored markers to track the angles and compared the real-world movement to what BoxTwin predicted in the computer. The result? BoxTwin was spot on. It didn't just get the movement right for a second; it kept getting it right over a long sequence, accurately tracking how the object slowly changed its shape and got "tired" from being folded again and again.

In the second, more impressive test, they used a dual-armed robot (called Trossen Aloha) to grab, fold, and move these tricky objects. They recorded every move the robot made and then replayed those exact moves in the BoxTwin simulation. The digital twin predicted the robot's actions and the object's reaction with high accuracy, even for complex objects with multiple joints. The simulation matched the real world so well that the team believes robots can now use these digital twins to plan their moves safely without needing to crash into things or fail repeatedly in the real world.

In short, BoxTwin is a leap forward because it stops pretending that bendy, breakable objects are simple. By combining video learning with a deep understanding of how materials bend, stay bent, and break, it gives robots a way to anticipate the messy, unpredictable nature of the real world. This means robots might soon be able to handle delicate tasks—like folding laundry or assembling flimsy parts—with the same confidence they currently have when moving heavy, solid boxes.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →