EvoPlan: Evolutionary Neuro-Symbolic Robot Planning with Spatio-Temporal Guarantees
This paper introduces EvoPlan, a neuro-symbolic framework that integrates evolutionary search to mine spatio-temporal constraints from demonstrations and combines them with an evolutionary PDDL planner and constrained execution loop to enable LLM-based robots to generate safe, executable plans with formal guarantees in open-world environments.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a robot to navigate a busy world, like a delivery bot in a factory or a self-driving car in a city. You have two main tools to help it, but both have a major flaw:
- The "Creative Writer" (AI/LLM): This is a smart AI that can read your instructions ("Go to the cafeteria") and come up with a clever plan. It's great at understanding context and fixing mistakes. The problem: It's a bit reckless. It might come up with a plan that sounds good but is physically impossible, or it might accidentally drive through a red light because it didn't "think" about the rules.
- The "Strict Accountant" (Classical Planners): This is a rigid, rule-based system. It only works if you give it a perfect, mathematical description of the world. The problem: It's terrible at understanding natural language or fixing its own mistakes. If the world changes slightly, it freezes.
EvoPlan is a new framework that combines the best of both worlds. It uses the "Creative Writer" to dream up plans and the "Strict Accountant" to make sure they are safe and executable. Here is how it works, broken down into three simple steps:
1. The "Shadow Coach" (Learning the Rules)
Before the robot ever moves, the system needs to learn the "unwritten rules" of the road or the factory floor.
- The Problem: The researchers only have videos of good behavior (experts driving safely or people walking politely). They don't have videos of bad behavior to show the robot what not to do.
- The Solution: The system acts like a creative writing coach. It takes a video of a good drive and asks the AI: "What if the driver sped up? What if they ran a red light?" It invents these "bad" scenarios (counterfactuals) to create a contrast.
- The Result: By comparing the "good" videos against the invented "bad" ones, the system learns a single, universal rulebook (called an STL constraint). Think of this as a "Shadow Coach" that whispers in the robot's ear: "Never go faster than X," or "Always stay Y meters away from people." This rulebook applies to every move the robot makes.
2. The "Evolutionary Editor" (Building the Plan)
Now, the robot needs a plan to get from Point A to Point B.
- The Process: The "Creative Writer" (AI) proposes a plan. But instead of just accepting it, the system puts it through a rigorous editing process.
- The Loop:
- The AI writes a draft plan.
- A "Validator" (the accountant) checks it. If the plan is broken (e.g., "You can't open a door that doesn't exist"), the validator marks the specific error.
- The AI reads the error, fixes that specific part, and tries again.
- This happens over and over, like a writer refining a story until it's perfect.
- The Magic: The system keeps the "good parts" of the plan that have already been checked and only rewrites the broken parts. Over time, the plan grows longer and more reliable until it is fully valid.
3. The "Safety Gatekeeper" (Real-Time Execution)
Finally, the robot starts moving. This is where the "Shadow Coach" and the "Safety Gatekeeper" take over.
- The Check: The robot doesn't just execute the plan blindly. Before it takes a single step (or turns a wheel), the system checks that specific movement against the rulebook learned in Step 1.
- The Safety Net:
- Scenario A: The robot plans to turn left. The Gatekeeper checks: "Is there a pedestrian there? Is the speed safe?" Yes. The robot moves.
- Scenario B: The robot plans to turn left. The Gatekeeper checks: "Wait, a pedestrian is too close!" No. The robot stops immediately.
- The Re-Plan: If the Gatekeeper says "No," the robot doesn't crash. It locks in all the safe steps it has already taken, updates its current location, and asks the "Evolutionary Editor" to write a new plan for the rest of the journey.
Why This Matters
The paper shows that this system works better than using just the AI or just the rules.
- In Driving: When tested on driving data, the system reduced collisions and red-light violations significantly without needing to retrain the driver AI. It just added a "guardrail" that stopped bad moves.
- In Navigation: When tested in complex, open-world puzzles (like finding objects in a house), the system solved more tasks than other AI planners, even when the words used in the instructions didn't perfectly match the robot's vocabulary.
In short: EvoPlan is a robot planning system that uses AI to be creative and flexible, but forces it to wear a "safety helmet" (learned from data) and a "rulebook" (checked by code) to ensure it never does anything dangerous or impossible. It's like having a brilliant but impulsive driver who is constantly guided by a very strict, very smart co-pilot.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.