← Latest papers
💻 computer science

DiscreteRTC: Discrete Diffusion Policies are Natural Asynchronous Executors

The paper proposes DiscreteRTC, a framework that leverages the native unmasking capability of discrete diffusion policies to enable efficient, fine-tuning-free asynchronous execution for physical AI, achieving significantly higher success rates and lower inference costs compared to existing flow-matching-based approaches.

Original authors: Pengcheng Wang, Kaiwen Hong, Chensheng Peng, Katherine Driggs-Campbell, Masayoshi Tomizuka, Chenfeng Xu, Chen Tang

Published 2026-04-29
📖 5 min read🧠 Deep dive

Original authors: Pengcheng Wang, Kaiwen Hong, Chensheng Peng, Katherine Driggs-Campbell, Masayoshi Tomizuka, Chenfeng Xu, Chen Tang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Problem: The "Thinking vs. Acting" Dilemma

Imagine you are playing a fast-paced video game, like a tennis match against a computer.

  • The Robot's Job: The robot needs to hit the ball.
  • The Problem: The robot's brain (the AI) takes a moment to calculate the perfect swing. While it is "thinking," the ball keeps flying.
  • The Old Way (Synchronous): The robot stops, thinks, calculates, and then swings. By the time it swings, the ball has already passed. This is fatal for dynamic tasks.
  • The Better Way (Asynchronous): The robot needs to think while it is already moving. It should keep swinging based on its last thought while simultaneously calculating the next swing.

The Current Solution: "Real-Time Chunking" (RTC)

To solve this, engineers use a method called Real-Time Chunking (RTC).

  • The Analogy: Imagine the robot is writing a story in chunks of 10 words at a time.
    1. It writes words 1–10.
    2. While it is writing words 11–20, it starts executing words 1–10.
    3. The Glitch: When the robot switches from the first chunk to the second, the transition can be jerky. It's like a writer suddenly changing their handwriting style mid-sentence. The robot might jerk its arm because the new plan doesn't perfectly match the old one.

To fix this jerkiness, the current best method (using Flow-Matching) tries to "paint over" the transition. It freezes the words it has already written and tries to "inpaint" (fill in) the rest of the sentence to make it smooth.

  • The Catch: This "inpainting" is like asking a painter to fix a painting while they are still mixing the paint. It requires a special, complicated guide (a heuristic) to tell the AI how to blend the old and new. It's slow, requires extra training, and often feels clunky.

The New Solution: DiscreteRTC

The authors propose a new way called DiscreteRTC. They swap the robot's "brain" (the policy) for a Discrete Diffusion Policy.

The Analogy: The "Fill-in-the-Blanks" Game
Imagine the robot isn't writing a story from scratch. Instead, it's playing a game of "Fill-in-the-Blanks" (like a Mad Libs or a crossword puzzle).

  • How it works: The robot is trained to look at a sentence with some words hidden (masked) and guess what the missing words are.
  • The Magic: Because the robot is already an expert at guessing missing words, it doesn't need to learn a special trick to "inpaint" the transition. It just does what it was born to do.

Here is why this new method is a game-changer, according to the paper:

1. No Extra Training Needed (Fine-Tuning Free)

  • Old Way: You had to teach the robot a special "transition skill" after it was already trained, which was hard and risky.
  • New Way: The robot is already trained to fill in blanks. When it needs to switch chunks, it just treats the transition as a new "fill-in-the-blanks" puzzle. It works perfectly right out of the box.

2. Natural Guidance (No Complicated Rules)

  • Old Way: Engineers had to write complex, manual rules (heuristics) to tell the AI how to blend the old and new actions. It was like giving a driver a manual on how to steer, rather than letting them drive.
  • New Way: The robot naturally knows how to handle the transition. If it has already figured out the first few words of the new chunk, it just stops there and waits for the next thought. The "unmasked" words naturally guide the next step without needing a manual.

3. Faster and Cheaper (Lower Inference Cost)

  • Old Way: The "painting over" method required the robot to do double the work (calculating the original plan plus the correction), making it slower.
  • New Way: Since the robot only needs to "unmask" (reveal) the words it hasn't figured out yet, it skips the work it has already done. It's like reading a book where you only read the pages you haven't seen yet, rather than re-reading the whole book every time you turn a page.
    • Result: It is about 30% faster (0.7x computation) than the old method.

The Results: Does it actually work?

The authors tested this in two ways:

  1. Simulated World (Kinetix): A digital playground with moving targets.
    • Result: The new method solved tasks more often and handled delays much better than the old method.
  2. Real Robot (Real World): A physical robot arm trying to pick up moving objects and place them on moving platforms.
    • Result: The old method failed completely (0% success) on the hardest moving tasks. The new method succeeded 95–100% of the time.
    • Speed: The new method was also faster in real-time execution.

Summary in One Sentence

DiscreteRTC is a new way for robots to think while they move that treats action planning like a "fill-in-the-blanks" game, making it naturally smoother, faster, and requiring zero extra training compared to the current clunky methods.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →