MGRPO: Mamba-based Multi-Agent Group Relative Policy Optimization for Biomimetic Underwater Robots Pursuit
This paper proposes MGRPO, a novel framework combining Mamba-based policies with group relative policy optimization under a CTDE paradigm to enhance long-horizon decision-making, inter-agent coordination, and training stability for cooperative pursuit tasks in biomimetic underwater robots.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a school of robotic fish how to work together to catch a fast, slippery, and clever "prey" fish in a giant swimming pool. This is the challenge the researchers tackled in this paper.
Here is the story of their solution, M2GRPO, broken down into simple concepts and everyday analogies.
1. The Problem: The "Amnesia" and "Noise" of the Deep
Traditional robot teams are like a group of friends trying to catch a ball in a foggy room.
- The Fog (Partial Observability): The robots can't see the whole pool; they only see what's right in front of them.
- The Amnesia (Long-Horizon Decisions): If a robot forgets where the prey was 10 seconds ago, it can't predict where the prey will be next. Old robot brains (called MLPs) are like short-term memory; they forget the past quickly.
- The Noise (Coordination): If one robot panics, the whole team might get confused.
The researchers needed a system that could remember the past, talk to each other without speaking, and learn quickly without needing a supercomputer.
2. The Brain: "Mamba" (The Super-Short-Term Memory)
The researchers gave the robots a new type of brain called Mamba.
- The Analogy: Imagine a standard robot brain is like a person reading a book one page at a time, forgetting the first page by the time they get to the tenth.
- The Mamba Upgrade: Mamba is like a person who can read the whole book, remember every character's motivation, and predict the ending, all while running a marathon. It uses a "selective state-space" mechanism. Think of it as a smart filter that only keeps the important memories (like "the prey turned left") and forgets the noise (like "a bubble floated by").
- Why it matters: This allows the robots to plan far ahead. They don't just chase the prey's current position; they chase where the prey will be in the future.
3. The Teamwork: "Group Relative" Learning
Usually, to teach a team, you need a coach (a "Critic") standing on the sidelines telling everyone exactly how well they did. But in underwater robotics, having a coach is expensive and slow.
The researchers used a trick called Group Relative Policy Optimization (GRPO).
- The Analogy: Imagine a classroom where the teacher doesn't give grades based on an absolute test score. Instead, the teacher says, "Who did better than the average of the class today?"
- How it works: The robots run the same chase game 10 times at the same time (in parallel). At the end, they compare their scores. If Robot A caught the prey faster than the average of the group, it gets a "high five" (positive reward). If it was slower, it gets a "thumbs down."
- The Benefit: They don't need a complex coach to calculate the "perfect" score. They just compare themselves to each other. This makes training faster, cheaper, and more stable.
4. The Architecture: CTDE (The "Study Hall" vs. "The Exam")
The system uses a method called CTDE (Centralized Training, Decentralized Execution).
- Training (The Study Hall): During practice, all the robots share their notes. They see the whole pool and talk to each other to figure out the best strategy together.
- Execution (The Exam): When the real game starts, the robots are alone in the water. They can't talk to the coach or see the whole pool. They have to rely only on their own eyes and their memory (the Mamba brain) to make decisions.
- The Result: The robots learn a complex team strategy in the study hall, but they can execute it perfectly even when they are alone in the dark.
5. The Real-World Test: Robotic Sharks
The researchers didn't just simulate this on a computer; they built real robotic sharks.
- These sharks have tails that wiggle like real fish, making them quiet and agile.
- They put them in a 4x4 meter pool.
- The Result: The M2GRPO sharks were incredibly successful. They caught the "prey" 97% of the time (compared to 80-90% for older methods) and did it faster.
- The Strategy: In the videos, you can see the sharks working like a pack of wolves. They don't just run in a straight line; they cut off escape routes, corner the prey, and take turns chasing.
Summary: Why is this a Big Deal?
This paper is like giving a school of robotic fish a super-memory and a team spirit that doesn't require a central boss.
- Old Way: Robots are short-sighted, forgetful, and need a supercomputer to coordinate.
- New Way (M2GRPO): Robots remember the past, predict the future, learn by comparing themselves to teammates, and can operate efficiently even with limited power.
This technology could soon be used for real-world tasks like searching for lost divers, inspecting underwater pipelines, or monitoring ocean health, where robots need to work together silently and intelligently in the deep, dark water.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.