← Latest papers
💻 computer science

CycleRL: Sim-to-Real Deep Reinforcement Learning for Robust Autonomous Bicycle Control

This paper presents CycleRL, a sim-to-real deep reinforcement learning framework utilizing Proximal Policy Optimization and domain randomization within NVIDIA Isaac Sim to achieve robust, high-performance autonomous bicycle control that successfully transfers from simulation to real-world hardware.

Original authors: Gelu Liu, Teng Wang, Zhijie Wu, Junliang Wu, Songyuan Li, Xiangwei Zhu

Published 2026-06-26
📖 5 min read🧠 Deep dive

Original authors: Gelu Liu, Teng Wang, Zhijie Wu, Junliang Wu, Songyuan Li, Xiangwei Zhu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine trying to teach a toddler to ride a bicycle. If you try to explain the complex physics of balance, friction, and momentum using a textbook, the toddler will likely get confused and fall. That is essentially the problem engineers face with traditional robot bicycles. They try to write perfect mathematical "textbooks" (models) for how the bike should behave, but the real world is messy, and those textbooks often fail when the wind blows or the road gets bumpy.

This paper introduces CycleRL, a new way to teach a robot bicycle to ride itself. Instead of giving the bike a textbook, the researchers let it learn by trial and error, just like a human child, but at superhuman speed.

Here is how they did it, broken down into simple concepts:

1. The "Virtual Playground" (Simulation)

You can't teach a robot to ride by letting it crash a real bike a million times; the bike would break, and the robot would be too slow to learn.

  • The Analogy: Think of this as a video game. The researchers built a super-realistic virtual world (using NVIDIA Isaac Sim) where they created a digital twin of a bicycle.
  • The Process: In this video game, the AI "child" tries to ride the bike. Every time it falls, it gets a "game over." Every time it stays upright and follows a path, it gets points.
  • The Learning: The AI uses a method called Deep Reinforcement Learning. It's like a dog learning tricks: if it does the right thing (stays balanced), it gets a treat (points). If it falls, it gets no treat. Over millions of attempts in the game, the AI figures out exactly how to twist the handlebars and shift its weight to stay upright.

2. The "Training Regimen" (Domain Randomization)

A major problem with video games is that what you learn in the game doesn't always work in real life. Maybe the virtual grass is too slippery, or the virtual wind is too strong.

  • The Analogy: Imagine training a soccer player only on a perfectly flat, dry field. When they go to a real match on a rainy, muddy field, they might slip.
  • The Solution: To fix this, the researchers used a technique called Domain Randomization. They didn't just let the AI practice on one perfect field. They made the virtual world chaotic and unpredictable:
    • Sometimes the bike was heavy; sometimes it was light.
    • Sometimes the tires were sticky; sometimes they were slippery.
    • Sometimes the wind blew hard; sometimes the road was bumpy.
    • Sometimes the bike started with a slight wobble.
  • The Result: By forcing the AI to learn how to ride in every possible weird scenario, the AI became so robust that when it finally stepped onto a real bike, it didn't care about the differences. It had already "seen it all" in the simulation.

3. The "Scorecard" (Reward Function)

How did the researchers tell the AI what "good" riding looked like? They created a complex scorecard with five different rules:

  1. Stay Alive: Don't fall over (Survival Reward).
  2. Go Fast: Match the speed you were told to go (Velocity Reward).
  3. Go Straight: Follow the direction you were told to go (Heading Reward).
  4. Don't Jerk: Move the handlebars smoothly, not wildly (Action Penalty).
  5. Be Efficient: Don't use too much energy (Action Magnitude Penalty).

The AI had to balance all these rules at once, learning that falling over was the worst outcome, but also that driving smoothly was better than driving erratically.

4. The Real-World Test (Sim-to-Real)

After the AI mastered the virtual world, they put it on a real, physical bicycle built in their lab.

  • The Hardware: They built a custom bike with a computer brain (NVIDIA Jetson), sensors to feel its balance (IMU), and motors to steer and drive.
  • The Result: The AI, which had never touched a real bike before, got on and started riding. It successfully balanced, followed paths, and handled different terrains like gravel and ramps.
  • The Stats: In the simulation, it stayed upright 99.9% of the time. On the real bike, it stayed upright 95% of the time. It could even recover from being pushed over, just like a skilled human rider.

Why This Matters

Traditional methods try to solve the bike problem with rigid math formulas. If the math is slightly wrong, the bike crashes.
CycleRL is different. It's like teaching a rider by letting them practice in a chaotic, ever-changing obstacle course until they become an expert. Because it learned through experience rather than rigid rules, it is much better at handling the unpredictable nature of the real world.

In short: The researchers taught a robot to ride a bike by letting it play a chaotic video game millions of times, making it so tough that it could ride a real bike perfectly on its first try.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →