← Latest papers
💻 computer science

LiDARDraft: Generating LiDAR Point Cloud from Versatile Inputs

LiDARDraft is a novel framework that generates realistic and diverse LiDAR point clouds from versatile inputs like text, images, and sketches by unifying them into 3D layouts to guide a rangemap-based ControlNet, thereby enabling controllable "simulation from scratch" for autonomous driving environments.

Original authors: Haiyun Wei, Fan Lu, Yunwei Zhu, Zehan Zheng, Weiyi Xue, Lin Shao, Xudong Zhang, Ya Wu, Guang Chen

Published 2026-07-30
📖 4 min read☕ Coffee break read

Original authors: Haiyun Wei, Fan Lu, Yunwei Zhu, Zehan Zheng, Weiyi Xue, Lin Shao, Xudong Zhang, Ya Wu, Guang Chen

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot how to drive a car. To do this safely, the robot needs to practice millions of miles in a virtual world before it ever touches a real steering wheel. But building these virtual worlds is incredibly hard and expensive. Usually, engineers have to manually place every tree, building, and other car in a 3D computer program, which takes forever. This is where a special kind of computer science called "generative AI" comes in. Think of it like a super-smart artist that has looked at millions of photos and can now draw new pictures that look real. Scientists have been trying to teach this artist to draw 3D maps of the world using laser scanners (called LiDAR), which are the "eyes" of self-driving cars that see in 3D. The big challenge has been that while the artist is great at drawing, it's terrible at following specific instructions. If you ask it to "draw a busy intersection," it might draw a forest instead, or put the cars in the sky. It's like trying to give a recipe to a chef who only speaks a different language; the result is often a mess.

This is the problem a team of researchers at Tongji University tackled with a new tool they call LiDARDraft. They wanted to find a way to let people create realistic 3D driving scenes using simple inputs like a sentence of text, a casual photo, or even a rough sketch, without needing to be 3D modeling experts. Their secret weapon is a clever middleman they call a "3D layout." Imagine you are building a city out of Lego bricks. Instead of trying to sculpt every single brick to look like a real car or tree, you just use simple blocks: a red box for a car, a green oval for a bush, and a flat gray plane for the road. This simple "Lego blueprint" is easy to make, but it holds all the important information about where things are and what they are. LiDARDraft uses this Lego blueprint as a bridge. It takes your simple input (like a text description), turns it into this Lego layout, and then uses a special guide (called ControlNet) to tell the AI artist exactly how to turn those simple blocks into a detailed, realistic 3D laser scan.

The researchers found that this approach works surprisingly well. By forcing the AI to follow the "Lego blueprint," they could generate high-quality 3D point clouds (the digital 3D maps) that matched the user's instructions perfectly. For example, if you typed "a car in front of me and a tree on the left," the system created a scene with exactly that arrangement, with the right shapes and positions. They tested this with text, images, and even existing 3D scans, and in every case, their method produced better results than previous attempts, which often added random, fake cars or got the road layout wrong. The team showed that this system could even take a single photo of a street and turn it into a full 3D driving simulation, or take a rough sketch and fill it in with realistic details.

What makes this particularly exciting is how flexible it is. The researchers demonstrated that you can edit the scene just by moving the "Lego blocks" in the layout. If you want to remove a car from the simulation, you just delete the red box from the blueprint, and the AI instantly regenerates the scene without that car, keeping everything else consistent. They also found that this method is efficient; it didn't need to be retrained from scratch every time, saving a massive amount of computer power. In their tests, the system could learn to follow these layout instructions in just 5,000 steps, which is a huge improvement over starting from zero. The results suggest that we are getting closer to a future where we can build entire self-driving training worlds just by describing them in plain English or showing a quick photo, making the process of teaching robots to drive much faster, cheaper, and safer.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →