DeCAL: Towards Physically-Grounded Dexterous Vision-Language-Action Models via Contact-Aware Latent Co-Imagination
DeCAL is a physically-grounded dexterous vision-language-action model that leverages a Mixture-of-Transformers architecture with adaptive visuo-tactile fusion and latent co-imagination to achieve state-of-the-art performance in contact-rich manipulation tasks by dynamically integrating tactile sensing and modeling physical dynamics.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Robots have long been masters of the factory floor, moving heavy objects along predictable paths with rigid precision. Yet, when we ask them to enter our homes and perform the delicate, messy tasks of daily life—screwing in a lightbulb, wiping a vase, or twisting a cap off a bottle—they often falter. The world is not a clean, static grid; it is a place of constant, subtle friction. To succeed, a machine must do more than just see; it must feel. It needs to understand the invisible physics of contact: the moment a finger slips, the pressure required to turn a knob, or the way a soft object deforms under touch. For decades, scientists have tried to bridge this gap by giving robots both eyes and skin, but teaching a machine to combine these senses into a single, fluid intuition has remained one of the most stubborn challenges in artificial intelligence.
A team of researchers has now taken a significant step forward with a new system called DeCAL. This is not merely a robot that can see and touch; it is a system designed to imagine the future of a physical interaction before it happens. By building a model that unifies understanding, imagination, and action, the researchers created a robot policy that can navigate the complex, contact-rich world of human manipulation with a level of dexterity previously unseen. The system was tested on six distinct tasks, ranging from wiping a vase to assembling small parts, and it consistently outperformed the most advanced existing methods. In these real-world trials, DeCAL achieved a success rate of up to 100 percent on simpler tasks and maintained a robust average success rate of 71 percent across all six challenges, even when the environment changed in unexpected ways.
The core problem the researchers tackled is that vision alone is often blind to the most critical moments of manipulation. When a robot hand reaches for an object, the fingers often block the camera's view, creating a "blind spot" exactly when the robot needs to know the most. Furthermore, visual data cannot tell a robot how hard it is pushing or if an object is about to slip. While previous attempts to add touch sensors to robots often treated them as simple add-ons, DeCAL integrates tactile information as a fundamental part of its thinking process. The system uses high-resolution sensors on each fingertip that can detect the shape of an object's surface as it deforms under pressure, as well as the precise forces being applied. This allows the robot to "feel" the world in a way that is as rich and detailed as what it sees.
What makes DeCAL truly unique is how it processes this information. Instead of simply reacting to what it sees or feels in the present moment, the system is trained to imagine what will happen next. It operates with three distinct but connected capabilities. First, it understands the current situation by looking at the visual scene, reading a language instruction, and feeling the contact points. Second, it engages in a form of "co-imagination," where it predicts how the visual scene and the tactile sensations will evolve in the next few seconds. It essentially asks itself, "If I turn the cap now, how will the force change, and what will the object look like a moment later?" Finally, it uses this imagined future to decide on the precise movements needed to complete the task. This forward-looking approach allows the robot to plan its actions based on the physics of the interaction, rather than just guessing based on past patterns.
To make this work, the researchers developed a special method for deciding when to trust the sense of touch. In many tasks, a robot might be moving through the air where touch is irrelevant, and then suddenly make contact with an object. If the robot paid equal attention to touch signals at all times, it would be confused by the noise of the air or the lack of contact. DeCAL uses a smart gating mechanism that acts like a volume knob for the sense of touch. When the robot is not touching anything, the system lowers the volume on the tactile signals, relying mostly on vision. The moment contact is made, the system instantly turns up the volume, allowing the tactile data to guide the movement with high precision. This dynamic adjustment ensures that the robot uses the right sense at the right time, preventing the sensory inputs from conflicting with one another.
The system was put to the test in a series of rigorous experiments involving a pair of robotic arms equipped with twenty-two-fingered hands. The researchers tasked the robot with six different activities: wiping a vase, erasing a whiteboard, assembling parts, twisting a cap, pipetting liquid, and screwing in a lightbulb. These tasks were chosen because they all require fine-grained control and constant physical contact. The robot was compared against several other state-of-the-art models, including those that relied only on vision and others that used touch but lacked the ability to imagine future states. The results were clear: DeCAL succeeded in completing the tasks far more often than any of the other systems. For instance, on the task of screwing in a lightbulb, where vision is often blocked by the hand and the task requires precise torque, DeCAL succeeded 40 percent of the time, while the next best model managed only 35 percent. On the task of wiping a vase, DeCAL achieved a perfect 100 percent success rate, whereas the best competing vision-only model succeeded only 30 percent of the time.
Perhaps even more impressive was the system's ability to handle situations it had never seen before. The researchers tested DeCAL in environments with different backgrounds, cluttered with random objects, under dim lighting, and with entirely new objects that had different shapes and sizes. In these "out-of-distribution" scenarios, where the robot could not rely on memorized patterns, DeCAL continued to perform strongly. When asked to twist a cap on a new, unseen cup, the system succeeded 75 percent of the time, significantly outperforming its competitors. This suggests that the robot is not just memorizing a specific set of movements but is actually learning the underlying physical principles of how objects interact, allowing it to adapt to new challenges on the fly.
The researchers also looked closely at how the system learned to predict the future. By analyzing the robot's internal predictions, they found that it could accurately forecast the forces and deformations that would occur during an interaction. When the robot predicted how much force would be needed to turn a cap or how a soft surface would deform, these predictions aligned closely with the actual physical reality. This ability to simulate the physics of the world internally is what gives the robot its "common sense." It allows the system to understand that if it pushes too hard, the object might slip, or if it turns too slowly, the task will take too long. This implicit knowledge of physics, derived from the combination of sight and touch, is what enables the robot to act with such fluidity and confidence.
Despite these successes, the authors are careful to note the limitations of their work. The system relies heavily on the accuracy of its tactile sensors, and if those sensors become noisy or miscalibrated, the robot's performance could degrade. Additionally, the current setup requires a human operator to demonstrate the tasks using a teleoperation system, which does not provide the operator with a sense of touch. This means the training data might not perfectly capture the subtle adjustments a human would make if they could feel the object themselves. The researchers also point out that their model has not yet been trained on a massive scale of diverse tactile data, suggesting that future improvements could come from exposing the system to an even wider variety of physical interactions.
The work represents a shift in how we think about robotic intelligence. Rather than building robots that are simply faster or stronger, this approach focuses on building machines that can understand the physical world as a dynamic, interactive place. By giving the robot the ability to imagine the consequences of its actions and to adapt its senses to the situation, the researchers have created a system that feels less like a machine following a script and more like an agent capable of genuine interaction. As the technology matures, the hope is that these systems will eventually be able to perform the complex, contact-rich tasks that define so much of human life, from cooking and cleaning to caring for the elderly, with a dexterity that matches our own. The path forward involves refining these sensors, expanding the training data, and continuing to bridge the gap between seeing, feeling, and doing.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.