GROVE: Grounded Pedestrian Simulation via Natural Language for Interactive Social Robot Navigation
GROVE is a text-to-scenario pedestrian simulation framework that leverages natural language prompts to dynamically generate realistic, socially complex human behaviors across multiple simulation environments, thereby bridging the sim-to-real gap for training interactive social robot navigation.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a director trying to film a scene for a movie about a robot walking through a busy hospital. In the old days, to get the "actors" (the digital people) to behave correctly, you would have to manually tell each one exactly where to stand, when to walk, and who to bump into. It was like choreographing a dance where you had to write down every single step for hundreds of dancers. If you wanted to change the scene from a "calm morning" to a "rushed evacuation," you'd have to rewrite the entire script from scratch.
GROVE is a new tool that changes how we create these digital crowds. Instead of writing a script, you just talk to the computer. You say something like, "Make a busy hospital hallway where people are rushing to the exit because of an alarm," or "Create a line of people waiting at a pharmacy counter." GROVE then instantly builds a realistic simulation of that exact scene.
Here is how it works, broken down into simple parts:
1. The "Brain" and the "Librarian" (RAG and LLM)
Think of GROVE as having two main helpers:
- The Librarian (RAG): Before the computer tries to write the scene, it checks a massive library of pre-written "behavior notes." If you ask for a "queue," the librarian finds the specific notes on how people stand in line. If you ask for an "emergency," it finds notes on how people run toward exits. This stops the computer from making things up (hallucinating) or forgetting the rules of how people actually behave.
- The Director (LLM): Once the librarian hands over the right notes, the Director (a Large Language Model) writes the "script" for the scene. It doesn't just say "walk here"; it creates a Behavior Tree. Think of a Behavior Tree like a flowchart or a decision tree. It tells the digital people: "If you see a robot, stop. If you hear an alarm, run to the door. If you are in a line, wait your turn."
2. The "GPS" and the "Flow" (Global Planner & Velocity Fields)
Just giving people a script isn't enough; they need to know how to move without crashing into walls.
- The GPS (Global Planner): GROVE uses a smart map system (called Theta*) to draw invisible paths for the people. It ensures that even if the "script" says "go to the pharmacy," the digital people know how to walk around the furniture to get there without walking through walls.
- The Flow (Velocity Fields): In chaotic situations (like an emergency), GROVE doesn't just tell each person individually where to go. Instead, it creates an invisible "wind" or current (a velocity field) that pushes the whole crowd in the right direction, like water flowing down a river. This keeps the crowd moving together smoothly without them bumping into each other.
3. The "Modes" (Presets)
To make things even easier, GROVE has three special "modes" or presets, like different camera filters on a phone:
- Emergency Mode: The computer switches into "panic mode." It ignores small talk and focuses entirely on getting everyone to the exit as fast as possible, creating a realistic rush.
- Queuing Mode: The computer sets up an orderly line. It knows exactly how to space people out so they don't crowd the counter.
- Normal Mode: This is for everyday life. People might chat, walk in groups, or stop to look at things, just like in a real office or home.
Why is this a big deal?
The paper compares GROVE to other tools and finds that:
- Old tools often made people walk in straight lines through walls or fail to form realistic lines.
- GROVE creates scenes that look and act much more like real life. When tested, it scored higher on "realism" and "following instructions" than other methods.
- It works directly with popular robot simulation software (like Isaac Sim and Gazebo), meaning robot developers can use these realistic crowds to train their robots to be safer and more polite in the real world.
In short: GROVE turns a simple text sentence into a complex, realistic crowd simulation. It combines a smart "librarian" to find the right behaviors, a "director" to write the script, and a "GPS" to ensure everyone moves safely, all to help train robots to navigate our busy, human-filled world.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.