← Latest papers
💻 computer science

How Should a Simulation-to-Reality Transfer Budget Be Spent?

This paper argues that in simulation-to-reality transfer, allocating a measurement budget to identify specific system parameters is more effective than broadly randomizing dynamics, suggesting that pipelines should prioritize measuring identifiable parameters and reserve randomization only for remaining uncertainties.

Original authors: Syed Hamzah Rizvi, Yash Vardhan Tomar

Published 2026-06-23
📖 4 min read☕ Coffee break read

Original authors: Syed Hamzah Rizvi, Yash Vardhan Tomar

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot to swing a pendulum perfectly. You can't just let the robot practice on the real metal arm because it's expensive, slow, and might break. So, you teach it in a video game (a simulation) first.

But here's the problem: The video game isn't exactly like the real world. The real robot might be slightly heavier, or the arm slightly longer than the game thinks. If you teach the robot only in the game, it might fail when you put it on the real machine. This difference is called the "reality gap."

To fix this, engineers have two main tools, but they both cost the same precious resource: time on the real robot. You can't just magically get more real-world practice time. So, you have to decide how to spend your limited "real-robot minutes."

The paper asks: Should you spend your time measuring the robot to get exact numbers, or should you spend your time teaching the robot to handle a wide variety of "what-if" scenarios?

The Two Strategies

  1. System Identification (The "Measurer"): You use your real-robot time to take measurements. You find out the exact weight of the bob and the exact length of the rod. Then, you update your video game to match those exact numbers. You are trying to make the simulation a perfect twin of reality.
  2. Domain Randomization (The "Gambler"): Instead of trying to find the exact numbers, you use your real-robot time to justify teaching the robot in a video game where the weight and length change randomly every single time. You hope the robot learns a "super-skill" that works no matter what the numbers are.

The Experiment: A Controlled Test

The authors set up a clever experiment. They didn't use a real robot (which would take too long). Instead, they created a "hidden" simulation that acted as the "real world" and a "training" simulation for the robot. They knew the "real" numbers (a mass of 2.0 and a length of 1.5), but the robot didn't.

They gave the robot a fixed budget of "real-world" practice rolls. They then tested different ways to spend that budget:

  • Option A: Spend all rolls measuring the hidden numbers to get a single, precise estimate.
  • Option B: Spend some rolls measuring, then train the robot on a wide range of numbers around that estimate.
  • Option C: Spend zero rolls measuring and just train the robot on a huge, random range of numbers (hoping the truth is somewhere in there).

The Surprising Results

The paper found that measuring is almost always better than guessing.

Here is the breakdown using a simple analogy:

  • The "Measurer" Wins: Even a tiny bit of measurement (just 5 or 10 practice rolls) closed most of the gap between the game and reality. Once the robot knew the approximate weight and length, it performed incredibly well.
  • The "Gambler" Loses: Once the robot had any real data, trying to teach it to handle a wide range of random weights actually made it worse. It was like trying to teach a driver to handle every possible car in the world, when they just needed to learn how to drive their specific car.
  • The "Perfect Range" Myth: Even when the authors created a randomization range that was guaranteed to include the true weight and length, it still didn't work as well as simply measuring the robot first. A policy trained to handle "everything" ended up being mediocre at handling "anything specific."

The Bottom Line

If you have a limited amount of time to test a robot on real hardware, spend that time measuring the specific details of your robot first.

Don't try to teach the robot to be a generalist who can handle any random variation. Instead, use your real-world time to get the best possible estimate of your robot's actual physics, and then train your simulation to match that specific estimate.

The Catch (Limitations):
The authors are careful to note that this advice works best when the robot's physics are "identifiable"—meaning the simulation is capable of perfectly representing the real robot if you just get the numbers right. If the real robot has weird, unmodelable problems (like sticky friction or broken gears that the simulation can't represent at all), then this rule might change. But for standard, well-behaved robots, measure first, randomize later.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →