DeicticVLA: Unifying Instruction Modes Based on Language and Deictic Gestures in a Single VLA
DeicticVLA unifies language, vision-language, and visual instruction modes into a single Vision-Language-Action model by canonicalizing them into text prompts and deictic masks, demonstrating superior generalization to unseen objects and expressions compared to language-only baselines.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a robot arm sitting in a kitchen, ready to help. To tell it what to do, a human might say, "Pick up the red cup on the left." This seems simple, but for a machine, the world is often cluttered with many red cups, or cups that look nearly identical. If the human tries to be more specific, saying, "Pick up the third red cup from the left, the one next to the blue napkin," the robot might still get confused. Modern robots are becoming better at understanding language, but they often struggle when the instructions are too detailed or when the objects in front of them don't match the exact pictures they learned from during training. They might grab the wrong item simply because it looks slightly more familiar, ignoring the specific words the human just spoke.
To solve this, researchers have been exploring a different way to talk to machines: using a pointing gesture. Just as a human might say, "Pick up this one," while pointing a finger at the object, a robot can be taught to understand a click on a screen as a direct command. This approach, known as a deictic gesture, cuts through the confusion of language by showing the robot exactly what is meant. However, until now, robots have usually been trained to listen to words only, or to listen to words while being pointed at, but rarely to handle all these ways of communicating at once. A new study introduces a system that unifies these methods, allowing a single robot brain to switch seamlessly between listening to a sentence, listening to a sentence while being pointed at, or simply following a point without any words at all.
The researchers, working with a type of artificial intelligence called a Vision-Language-Action model, developed a system they call DeicticVLA. These models are like a robot's brain that has read millions of books and seen millions of images, learning how to connect what it sees with what it should do. The challenge was to make this brain flexible enough to accept three different types of input: a text command, a text command combined with a pointing gesture, or just a pointing gesture. In the first mode, the user types a full sentence. In the second, the user types a sentence that includes a word like "this" or "there" and clicks on the screen to define what that word refers to. In the third, the user clicks on the object they want moved and the place they want it to go, saying nothing at all.
To make this work, the team created a translation step that turns all three of these different inputs into the same internal format. When a user clicks on a screen, the system uses a powerful image-analysis tool to turn that single click into a precise outline of the object, a mask that highlights exactly what the human is touching. This mask is then combined with a text prompt. If the user spoke, the text prompt is their words. If they only clicked, the system uses a default phrase like "follow the visual instruction." This way, the robot's brain always receives a consistent package: a text instruction and a visual highlight of the target. The researchers then tested different ways to feed this visual highlight into the robot's brain. They tried painting a box around the object on the camera image, fading out the rest of the scene to make the object stand out, or feeding the outline as a separate layer of information that the robot processes alongside the image.
The team first tested these ideas in a simulated environment, a digital world where robots can practice thousands of times without breaking anything. They found that the system could successfully learn to handle all three instruction modes at once. When the robot was asked to perform tasks it had never seen before, such as picking up an object from a new arrangement of furniture or identifying an object by a new spatial description, the pointing methods worked significantly better than words alone. In one test involving identical black bowls, where the robot had to pick the one next to a specific reference object, the language-only approach failed most of the time. However, when the user pointed to the correct bowl, the robot succeeded almost every time. The study showed that the two-stage training method was crucial: the robot first learned to follow text commands, and then it learned to combine those commands with pointing gestures. This order helped the robot keep its ability to understand language while gaining the new skill of following a point.
To see if this worked in the real world, the researchers built a physical setup with a robotic arm, a camera, and a tablet. They taught the robot to pick up blocks, place them in bowls, and organize stuffed toys. They tested the robot with instructions it had never heard before, such as asking it to pick up the "leftmost" block when it had only been trained on "rightmost" instructions, or asking it to move a specific type of toy it had never seen. In these difficult scenarios, the robot that relied only on language struggled, often guessing wrong or failing to act. But when the user pointed to the target, the robot's success rate soared. In a test with completely new categories of stuffed toys, the language-only robot succeeded only about one-sixth of the time. The pointing-based methods, however, achieved a perfect success rate. The robot did not need to know the name of the toy or the category it belonged to; it simply needed to see where the human pointed.
The results suggest that the future of human-robot interaction may not be about choosing between talking and pointing, but about having both available at the same time. The system allows a user to start with a sentence, and if the robot seems confused, to simply point to clarify. Or, for a quick, routine task, the user can just point without saying a word. The research indicates that while language is powerful for describing complex ideas, pointing is a more reliable way to identify specific objects in a messy, changing world. By unifying these methods into a single system, the researchers have created a more flexible and robust way for humans to guide machines, ensuring that the robot understands not just what we say, but what we mean.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.