← Latest papers
🤖 AI

Learning Implicit Constitutive Laws for Dynamic 3D Gaussian Splatting from Monocular Videos

The paper presents GCA (Gaussian Constitutive Alignment), a framework that learns implicit constitutive laws for dynamic 3D Gaussian Splatting from monocular videos by combining LoRA-based adaptation with rank-based depth-geometric anchors and a constitutive prior regularizer to overcome the limitations of existing methods in physical interpretability and monocular stability.

Original authors: Xiaoyang Liu, Kai Han

Published 2026-08-25
📖 4 min read☕ Coffee break read

Original authors: Xiaoyang Liu, Kai Han

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

To understand how objects move and change shape, scientists have long tried to teach computers the hidden rules of physics. When a ball bounces, a rubber band stretches, or clay squishes under pressure, they are all following specific laws of material behavior. These laws, known as constitutive laws, describe how a substance reacts to forces. For a computer to create a realistic digital twin of a real-world object, it must learn these rules. The challenge has always been that while humans can easily guess these properties just by watching a video, computers struggle. They often get confused by the lack of depth in a flat image or fail to distinguish between different types of materials when the data is noisy. Without a clear way to learn these rules from a single camera view, digital models often look stiff or behave in ways that defy reality.

A team of researchers at the University of Hong Kong has developed a new method to solve this problem, allowing computers to learn these hidden physical laws simply by watching a video of a moving object. Their approach, called GCA, focuses on objects represented by thousands of tiny, glowing dots that form a 3D shape. The researchers start with a static 3D model of an object, created from multiple camera angles, and then feed the computer a single video of that same object moving and deforming. The goal is to teach the computer the specific rules that govern how that object bends, stretches, or squishes, without needing any other sensors or measurements.

The researchers found that previous methods often failed in this setting because they relied too heavily on matching colors or exact pixel positions, which can be misleading when lighting changes or the object rotates quickly. Instead, their new system uses two clever strategies to guide the learning process. The first strategy focuses on geometry. Rather than trying to match every single color in the video, the system looks at the relative depth of the object's surface. It checks if the order of points from near to far remains consistent, even if the exact distance is hard to tell. This allows the computer to build a stable understanding of the object's shape as it moves, ignoring the confusing noise that often plagues single-camera recordings.

The second strategy addresses the mystery of what the object is made of. Since the computer does not know if it is watching rubber, plastic, or fluid, the system tests several common physical models at once. It treats these known models not as rigid rules, but as gentle suggestions. The computer tries to fit the video data to these suggestions and sees which ones stay stable. If the true material is something unusual that doesn't fit any standard model, the system still works because it only uses these models as a guide to keep the learning process from going off track. This allows the computer to learn the specific behavior of the object without being forced into a box that doesn't fit.

In their tests, the researchers showed that this method works significantly better than existing techniques. On synthetic benchmarks where the true physics were known, their system reduced the error in predicting the object's shape by nearly half compared to the strongest previous method. The system remained robust even when the video data was sparse or noisy, a situation where other methods often failed completely. By combining a flexible learning approach with these two guiding strategies, the researchers have created a way for computers to infer the invisible rules of material science just by watching a video, bringing us closer to digital worlds that move and react with the same natural logic as our own.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →