← Latest papers
💻 computer science

Flex-π\pi: A Multi-Stream World-Action Model with Compute Flexibility

Flex-π\pi is a 6B-parameter world-action model that leverages a frozen video-generation VAE to implicitly encode 3D geometry and object semantics alongside RGB, enabling a flexible, multi-stream training approach that achieves superior generalization and efficiency in real-world bimanual manipulation tasks without requiring new sensors or pre-training.

Original authors: Ge Yan, Jinghao Liu, Yuzhi Fan, Lei Cai, Minwen Liao, Jesse Zhang, Dieter Fox

Published 2026-08-12
📖 4 min read☕ Coffee break read

Original authors: Ge Yan, Jinghao Liu, Yuzhi Fan, Lei Cai, Minwen Liao, Jesse Zhang, Dieter Fox

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a robot trying to learn how to fold a shirt or fix a broken toy. For a long time, scientists taught robots by showing them thousands of videos of humans doing the task, hoping the robot would just "copy" the movements. But this is like trying to learn to drive a car by only watching the dashboard lights; you see the speedometer go up, but you don't understand the road, the curves, or the other cars. A newer, smarter approach called "World-Action Models" tries to fix this. Instead of just copying, these models try to predict what will happen next in the future. They ask, "If I move my arm this way, what will the world look like a second from now?" By guessing the future, the robot learns the rules of physics and how objects interact, making it much better at figuring out new tasks without needing a million examples. However, most of these smart robots only look at flat, 2D pictures (like a standard camera), which can be confusing when trying to grab a 3D object that might be hidden behind something else.

Enter FLEX-π, a new robot brain that solves this puzzle with a clever trick. Think of FLEX-π as a master chef who can taste a soup and instantly know exactly what spices are in it, how hot it is, and how thick it is, even though they only have a spoonful of the liquid. In the past, to get a robot to understand 3D shapes (like depth) or the meaning of objects (like "this is a cup, not a ball"), scientists had to give the robot expensive 3D cameras or teach it entirely new ways of seeing the world. FLEX-π, however, discovered a "free lunch." It realized that the same brain it uses to understand flat pictures could also understand 3D shapes and object meanings almost perfectly, without any extra training or new hardware.

The researchers built a 6-billion-parameter model (a very large digital brain) that learns to predict three things at once: what the next picture will look like, what the 3D shape of the scene will be, and what the important objects in the scene are. The magic happens because they trained the robot to imagine these three things together. If the robot sees a picture of a cup, it doesn't just see a flat image; it simultaneously imagines the cup's 3D depth and knows it's a "cup." Even better, they taught the robot to be flexible. During training, they sometimes hid the 3D data or the object names, forcing the robot to guess them using only the flat picture. This means that when the robot is actually working, it can run in "fast mode" using just the camera (ignoring the 3D and object data) and still be incredibly smart, or it can run in "super mode" using all the data if it has the time.

The results are impressive. When tested on real robots doing tricky tasks—like using a screwdriver to fix its own gripper with millimeter precision, or zipping up a soft, squishy pencil case—FLEX-π beat the best existing robot brains by a huge margin, sometimes performing 2 to 7 times better. It learned these complex skills with very few demonstrations, meaning it didn't need to watch humans do the task hundreds of times. Perhaps most surprisingly, even when FLEX-π was set to "fast mode" and only used the camera (ignoring the extra 3D and object data it learned), it was still faster and more successful than other top robots. The paper shows that by teaching a robot to imagine the future in 3D and with object meaning, we can make it much smarter and more adaptable, all without needing new sensors or slowing it down.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →