CRAFT: Coaching Reinforcement Learning Autonomously using Foundation Models for Multi-Robot Coordination Tasks
The paper proposes CRAFT, a framework that leverages Large Language Models to autonomously decompose complex multi-robot coordination tasks into subtasks and Vision Language Models to refine reward functions, thereby enabling effective learning and real-world transfer of coordination behaviors without manual reward design.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a team of two dogs to perform a complex dance routine, like walking through a narrow gate without bumping into each other, or lifting a heavy pot together without spilling the soup.
If you just throw them into the room and say, "Do it!" they will likely get confused, trip over each other, and never learn. This is the problem researchers at UC Berkeley faced with robots. They wanted robots to work together, but standard computer training methods were failing because the tasks were too long, too complicated, and the robots couldn't "talk" to each other to plan ahead.
Enter CRAFT.
Think of CRAFT as a super-smart, AI-powered coach that doesn't just watch the robots; it actively teaches them how to learn. Here is how this "coach" works, broken down into simple steps:
1. The Coach Breaks the Big Job into Small Drills (Curriculum Generation)
Imagine a soccer coach. They don't tell a new team, "Go win the World Cup!" immediately. Instead, they break it down: "Today, we practice passing. Tomorrow, we practice shooting. Next week, we practice defending."
CRAFT uses a Large Language Model (LLM)—a type of AI that is very good at understanding language and logic—to act as this coach. When given a big, scary task (like "Two robots must cross a gate"), the coach breaks it down into a list of smaller, easier steps:
- Step 1: Just learn to walk without hitting the gate.
- Step 2: Learn to walk past the gate.
- Step 3: Learn to coordinate so one waits while the other goes through.
2. The Coach Writes the Rules of the Game (Reward Generation)
In robot training, robots learn by getting "points" (rewards) for doing good things and losing points for bad things. Usually, humans have to write these rules, which is hard.
CRAFT's coach uses the AI to write the scoring rules for each small drill.
- Example: For the "walk past the gate" drill, the AI writes a rule that says, "You get points for getting closer to the gate, but you lose points if you hit it."
- The AI writes this rule in computer code that the robots can actually understand.
3. The Coach Watches and Critiques (Policy Evaluation)
Once the robots try the drill, a Vision-Language Model (VLM)—an AI that can "see" video and understand what's happening—watches the robots. It acts like a referee.
- Did they succeed? If yes, the coach says, "Great job! Let's move to the next drill."
- Did they fail? If they crashed or got stuck, the referee doesn't just say "Fail." It looks at the video and the score history to figure out why. Maybe the robots were trying too hard to balance and forgot to move forward.
4. The Coach Tweaks the Rules (Reward Refinement)
This is the magic part. If the robots fail, the coach doesn't give up.
- The "Referee AI" gives advice: "The robots are getting too many points for balancing and not enough for moving forward. The rule for moving needs to be stronger."
- The "Coach AI" takes this advice and rewrites the scoring rules to fix the problem.
- The robots try again with the new rules. They keep doing this loop (Try -> Watch -> Fix Rules -> Try Again) until they master that specific small drill.
The Result: From Simulation to Reality
The researchers tested this on two main challenges:
- Quadruped Navigation: Two four-legged robots (like dogs) learning to walk through a narrow gate without colliding.
- Bimanual Manipulation: Two robot arms learning to lift a heavy pot together while keeping it level.
What happened?
- Without this coach, other methods failed or learned weird, unnatural movements (like twisting their joints in impossible ways just to get points).
- With CRAFT, the robots learned to coordinate smoothly. They learned the basics first, then built up to the hard stuff.
- The Real-World Test: The best part? The researchers took the robots trained in the computer simulation and put them on real physical robots (Unitree Go1 and Go2). Without any extra tuning, the robots successfully walked through the gate in the real world, proving the training worked outside the computer.
In Summary
CRAFT is like a patient, intelligent tutor for robots. Instead of hoping robots figure out complex teamwork on their own, it:
- Simplifies the big task into small steps.
- Creates custom rules for each step.
- Watches the robots and fixes the rules when they make mistakes.
- Guides them from simple drills to complex coordination, eventually allowing them to perform real-world tasks they couldn't learn any other way.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.