← Latest papers
💻 computer science

Curriculum Reinforcement Learning for Quadrotor Racing with Random Obstacles

This paper proposes a novel vision-based curriculum reinforcement learning framework that combines multi-stage curriculum learning, domain randomization, and multi-scene updating to enable robust, high-speed quadrotor racing through unseen random obstacles, outperforming existing methods in both simulation and real-world experiments.

Original authors: Fangyu Sun, Fanxing Li, Yu Hu, Linzuo Zhang, Yueqian Liu, Wenxian Yu, Danping Zou

Published 2026-03-02
📖 5 min read🧠 Deep dive

Original authors: Fangyu Sun, Fanxing Li, Yu Hu, Linzuo Zhang, Yueqian Liu, Wenxian Yu, Danping Zou

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a toddler to ride a bicycle.

If you put them on a bike in a busy city with cars, pedestrians, and potholes immediately, they will likely crash. But if you start them in an empty parking lot, then move them to a quiet cul-de-sac, and finally take them to a park with a few slow-moving people, they will eventually learn to ride safely and fast.

This paper is about teaching a tiny, flying robot (a quadrotor drone) to race through a chaotic, obstacle-filled course at high speeds, using the same "step-by-step" logic.

Here is the breakdown of their solution, explained simply:

1. The Problem: The "Double-Edged Sword"

In drone racing, the goal is to fly as fast as possible through a series of hoops (gates).

  • The Conflict: To fly fast, the drone needs to go straight and aggressive. But to avoid crashing into random obstacles (like trees or poles), it needs to be cautious and slow down.
  • The Twist: The hoops themselves look like obstacles to the drone's camera. If the drone is too scared of obstacles, it will fly around the hoops instead of through them. If it's too aggressive, it will smash into the walls.
  • The Old Way: Previous AI methods were like students who memorized one specific test. If the test changed even slightly (new obstacle placement), they failed. They couldn't handle the "real world."

2. The Solution: "Curriculum Learning" (The School System)

The authors didn't just throw the drone into the deep end. They built a training school with three levels of difficulty:

  • Level 1 (The Playground): The drone learns to fly through empty hoops. No obstacles. It learns the basics of speed and steering.
  • Level 2 (The Obstacle Course): They add a few random obstacles, but the drone flies slowly. It learns to dodge things without panicking.
  • Level 3 (The Pro League): They remove the speed limits and fill the course with dense, random obstacles. The drone now has to combine everything: fly fast, hit the hoops, and dodge the chaos.

This is called Curriculum Reinforcement Learning. Just like a human athlete trains on a treadmill before running a marathon, the drone trains on easy tasks before tackling the hard ones.

3. The Secret Sauce: "Multi-Scene" Training

Imagine a teacher trying to teach 100 students.

  • Old Method: The teacher shows the same math problem to all 100 students at the same time. If the students memorize that one problem, they fail the test.
  • This Paper's Method: The teacher splits the 100 students into groups. Group A gets a hard problem, Group B gets a medium one, and Group C gets an easy one. They all learn different things and share what they learn.

The researchers used a "Multi-scene updating" strategy. They trained the AI on many different track layouts simultaneously. This prevented the drone from "cheating" by memorizing one specific track. Instead, it learned the concept of racing, making it adaptable to any new track it sees.

4. The "Reward System" (The Scoreboard)

In AI training, the drone gets "points" (rewards) for good behavior and loses points for bad behavior. The tricky part was balancing the score:

  • The Trap: If you only give points for avoiding obstacles, the drone will just hover in a safe corner and never race. If you only give points for speed, it crashes.
  • The Fix: The authors designed a complex scoreboard.
    • Points for: Moving forward, hitting the center of the hoop, and keeping a smooth flight path.
    • Penalties for: Hitting a wall, flying too fast for the current situation, or making jerky movements.
    • The Magic: They specifically tuned the points so the drone realized: "I can fly fast, but I must swerve slightly to avoid that pole, then zoom back to the hoop."

5. The Result: From Simulation to Reality

They trained the drone entirely in a computer simulation (a video game world) using a camera that sees depth (like human eyes).

  • The "Zero-Shot" Transfer: Usually, AI trained in a game crashes when put in the real world because the physics are different. But because they used "Domain Randomization" (changing the lighting, the drone's weight, and the obstacle positions randomly during training), the drone learned to be robust.
  • The Real-World Test: They put the AI on a tiny computer (a Raspberry Pi) strapped to a real drone.
    • Speed: It flew at 8 meters per second (about 18 mph).
    • Success Rate: It completed 100% of the races without crashing, even when obstacles were placed in random, unseen spots.
    • Comparison: It was significantly faster and more reliable than previous methods.

The Big Picture

Think of this paper as teaching a robot to be a Formula 1 driver who is also a ninja.

  • The Formula 1 part: It needs to be incredibly fast and precise.
  • The Ninja part: It needs to dodge flying shurikens (obstacles) without looking at a map.

By using a step-by-step training curriculum and a smart reward system, the authors taught the drone to be both. It's a major step toward drones that can autonomously race, deliver packages, or perform search-and-rescue missions in messy, unpredictable environments without human help.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →