ACME: A Multi-Cultural, Multi-Embodiment Social-Navigation Dataset
The paper introduces ACME, a large-scale, multi-cultural, and multi-embodiment dataset collected across five countries to advance social robot navigation research by providing diverse, goal-driven interaction data that addresses the lack of cultural and geographical diversity in existing resources.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a robot how to walk through a busy school hallway without bumping into anyone. This isn't just about avoiding walls; it's about understanding the invisible dance of human movement. This field is called social navigation. While a regular robot just needs to get from point A to point B without crashing, a social robot needs to know the unspoken rules of the crowd: when to step aside, how close is "too close," and which side of the hallway people prefer to walk on. These rules change depending on where you are in the world and what the robot looks like. A robot that is polite in one country might be rude in another, and a robot that looks like a dog might be treated differently than one that looks like a suitcase. Scientists care about this because as robots start working in our homes, hospitals, and streets, they need to fit in naturally without making us feel awkward or unsafe.
Enter ACME, a massive new project that acts like a giant, global training camp for these social robots. The researchers behind ACME realized that previous robot training data was too limited, like teaching a driver to drive only in a quiet, empty parking lot and then expecting them to handle a chaotic Tokyo intersection. To fix this, the team went on a world tour, collecting data from 8 different locations across 5 countries (including the US, Japan, Singapore, Germany, and Spain). They didn't just use one type of robot; they deployed 7 different robot "bodies" (or embodiments), ranging from four-legged dog-like robots to wheeled suitcase bots and tall service robots. They recorded over 29 hours of the robots moving through crowds and 43.5 hours of overhead video tracking how people walked around them.
The paper suggests that by feeding this diverse, messy, real-world data into computer models, we can teach robots to be much better at reading the room. The team found that robots trained on this new dataset face much harder challenges than those trained on older data, which is actually a good thing—it means the robots are learning to handle the real, unpredictable world. They also discovered that culture matters a lot: people in some countries tend to walk on the right, while others prefer the left, and the way people react to a robot depends heavily on whether the robot looks like a friendly dog or a rolling box. Furthermore, the study shows that robots using speech (like saying "Excuse me") interact differently with crowds depending on their size and shape. While the paper doesn't claim to have solved the problem of perfect robot navigation, it suggests that ACME provides the most comprehensive "textbook" yet for teaching robots how to navigate the complex, cultural, and physical dance of human spaces.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.