ReLI: Cross-Lingual Language-to-Action Grounding for Human-Robot Interaction
ReLI is a cross-lingual framework that adapts large-scale foundation models to enable autonomous agents to understand natural language instructions, reason semantically, and execute tasks effectively across over 140 diverse languages, including low-resource and creole varieties, thereby overcoming the limitations of existing high-resource language-only systems.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a world where a robot can understand a command given in any language, from the most widely spoken tongues to rare dialects spoken by small communities. For decades, the dream of natural human-robot collaboration has been held back by a simple barrier: language. While modern robots are becoming increasingly capable of seeing their surroundings and moving through them, they have largely been trained to understand only a handful of major languages, such as English or Mandarin. This limitation excludes billions of potential users who speak other languages, dialects, or creoles, effectively locking them out of the future of automation. The field of human-robot interaction relies on the ability to translate a spoken or written request into a physical action, a process known as grounding. Until now, this translation has been a one-way street, working well only for those with access to high-resource languages, leaving the rest of the world behind.
A team of researchers has developed a new system called ReLI, designed to break down these linguistic walls. ReLI allows autonomous agents to converse naturally, reason about their environment, and execute tasks regardless of the language in which the instruction is given. The system does not rely on translating the user's speech into English first, nor does it require the robot to be retrained for every new language. Instead, it leverages the inherent ability of large-scale pre-trained computer models to understand the meaning behind words across hundreds of different languages simultaneously. By combining these language models with visual sensors that allow the robot to "see" objects and spaces, ReLI can take a command like "find the red chair" spoken in a rare dialect and turn it into a precise physical movement, all without the user needing to know any technical code or speak a dominant global language.
The researchers tested this system in both simulated environments and real-world settings using wheeled robots and four-legged machines equipped with cameras and depth sensors. They challenged the robots with over 70,000 multi-turn conversations in more than 140 different languages. These languages were carefully selected to represent a wide spectrum of the world's linguistic diversity, including high-resource languages with vast digital data, low-resource languages with limited data, and vulnerable languages or creoles that are often ignored by technology. The tasks ranged from simple, single-step commands, such as moving forward a specific distance, to complex, multi-step missions that required the robot to navigate through a room, identify an object based on a description, and report back what it found.
The results showed that the system performed with remarkable consistency across this vast linguistic landscape. In nearly all the languages tested, the robot correctly understood the user's intent and executed the task with a success rate exceeding 83 percent. For many of the most common languages, the success rate climbed above 97 percent. Even for the most vulnerable and under-resourced languages, the system maintained high levels of accuracy, successfully parsing instructions and guiding the robot to its goal. The time it took for the robot to process a command and begin moving remained stable, averaging between 2.1 and 2.6 seconds, regardless of whether the instruction was given in a major world language or a rare dialect. This speed and reliability suggest that the system is not just a theoretical concept but a practical tool capable of functioning in real-time.
To ensure the technology was truly inclusive, the researchers also conducted a study with human participants who interacted with the robots in their native languages. Thirty-four raters, fluent in 13 different languages, tested the system in a real-world laboratory setting. They reported that the interactions felt natural and responsive, with the vast majority noting that they did not perceive any difference in performance based on the language they used. This finding is crucial, as it indicates that the system does not favor one language over another, nor does it create a hidden gap where users of certain languages receive a poorer experience. The system successfully handled code-switching, where a user might mix words from two different languages in a single sentence, and it managed complex spatial reasoning, such as navigating to a location described only by its function, like "the place where food is cooked."
The core of this achievement lies in how the system connects language to action. When a user speaks or types a command, the system first normalizes the input, converting speech to text if necessary, and then passes it to a large language model that acts as the robot's brain. This model interprets the intent of the command, breaking it down into a sequence of logical steps. It then consults the robot's visual sensors to identify the objects and spaces mentioned in the command. If a user asks the robot to go to a specific chair, the system uses computer vision to locate that chair in the room, determine its exact position in three-dimensional space, and calculate the path the robot needs to take to reach it. Finally, the system generates a plan and, for safety, often asks for user confirmation before executing complex or long-distance movements. This entire process happens seamlessly, bridging the gap between human thought and machine action without the need for explicit translation or specialized training for each new language.
The study also explored the limits of this technology, revealing that while the system is robust, it is not infallible. The most significant challenges arose when the robot had to identify objects that looked very similar to one another or when the visual environment was cluttered, making it difficult to distinguish the target. Additionally, tasks that required navigating through complex, multi-step sequences were slightly more prone to error than simple, direct commands, simply because there were more opportunities for a mistake to occur in the chain of reasoning. However, even in these difficult scenarios, the system's performance remained strong, and the researchers noted that the errors were often related to the physical limitations of the robot's sensors or the complexity of the environment rather than a failure to understand the language.
By demonstrating that a single framework can handle over 140 languages with high accuracy, this work suggests a new path forward for human-robot interaction. It moves away from the idea that robots must be taught a specific language to be useful, and toward a model where the robot adapts to the user's native tongue. The researchers released their data and code to the public, inviting others to build upon this foundation. The ultimate goal is not just to make robots that can speak many languages, but to create a future where the ability to command a machine is a universal right, accessible to anyone, anywhere, in the language they know best. This shift promises to democratize access to automation, ensuring that the benefits of robotics are shared by all of humanity, not just those who speak the languages of the digital elite.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.