A Numerically-Robust ROS 2 Port of iG-LIO: Diagnosing and Fixing Toolchain-Induced Failures in Incremental GICP LiDAR-Inertial Odometry
This paper presents a numerically robust ROS 2 Jazzy port of the iG-LIO LiDAR-inertial odometry system, detailing the diagnosis and resolution of critical toolchain-induced failures—specifically QoS mismatches and uninitialized parallel-reduce accumulators—while adding support for modern Ouster, Velodyne, and Livox sensors.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a super-smart robot explorer named iG-LIO. This robot is a master navigator; it combines a spinning laser scanner (LiDAR) and a motion sensor (IMU) to build a perfect 3D map of the world while figuring out exactly where it is. The original version of this robot was built for an old operating system called ROS 1.
Recently, a team of engineers tried to move this robot to a brand-new, modern operating system called ROS 2. They thought, "It's just a translation job! We'll keep the robot's brain exactly the same, just change the language it speaks." They did the translation, and the robot started up. But then, disaster struck: the robot's brain started screaming nonsense, filling its memory with "NaN" (Not a Number) errors and crashing. It was like a perfectly healthy car that, after a paint job, suddenly refused to drive because the new gas station pump didn't fit the nozzle.
The team realized the robot's brain (the math) was fine. The problem was the environment it was now living in. They found two sneaky culprits hiding in the new operating system that broke the robot, and they fixed them.
The First Culprit: The "Best Effort" Mix-Up
Imagine the robot's motion sensor (the IMU) is a frantic messenger running to the robot's brain, shouting updates about how the robot is tilting and spinning. In the old system, the robot's brain would wait patiently for every single message, no matter how crowded the hallway got.
In the new system, the robot was told to use a "Best Effort" delivery service. This is like a mail carrier who says, "I'll try to deliver these letters, but if the bag gets too full, I'll just drop the oldest ones and hope you get the rest." Because the robot was processing data slowly, the messenger got backed up. The "Best Effort" carrier started dropping and scrambling the order of the motion updates.
The robot's brain, which relies on a perfect, unbroken chain of motion data to stay balanced, got confused by the missing pieces. It tried to calculate a path based on a broken timeline and ended up with a math disaster (NaN values).
The Fix: The team changed the delivery contract. They told the robot's brain, "No more 'Best Effort.' We need Reliable delivery." They set up a massive waiting room (a queue of 2000 samples) so the messenger could dump all the updates without dropping a single one. They also added a safety guard: if the time between updates is weird (less than 0 seconds or more than 0.5 seconds), the robot just ignores that step instead of crashing.
The Second Culprit: The "Empty Box" Trap
The second problem was even sneakier. The robot's brain uses a super-fast parallel processing tool (called oneTBB) to do heavy lifting. Imagine a team of workers (threads) trying to count a pile of rocks. They split the pile, each worker counts their own stack, and then they add their totals together.
In the old system, the workers started with empty buckets that were magically zeroed out. In the new system, the workers were given buckets that looked empty but actually had random, dusty junk inside them because the new factory didn't clean them first. When the workers added their totals, they accidentally added this random junk to the final count. This "junk" was so bad it turned the robot's math into garbage (NaNs).
The Fix: The team didn't stop using the fast parallel workers. Instead, they wrapped the buckets in a special "Zero-First" sleeve. Now, before any worker starts counting, they are forced to wipe their bucket clean and start with exactly zero. This kept the speed of the parallel processing but ensured the math was clean.
New Gadgets and Better Maps
Beyond fixing the crashes, the team upgraded the robot's toolkit:
- New Scanners: They updated the robot to understand the newest laser scanners (like the Ouster OS0 and OS1 Rev 7) so it doesn't get confused by their new data formats. They also added support for a specific Velodyne Velarray M1600.
- Livox Flexibility: For Livox sensors, the robot can now work in two ways. It can talk to the special driver if you have it, or it can just listen to the standard data stream (like a Mid-360 sensor) without needing any extra software. This means users don't need to hunt down specific drivers anymore.
- Easy Settings: Everything is now controlled by a simple text file (YAML). You can tell the robot how reliable it needs to be, what to name its maps, and where to save its travel logs.
Did it Work?
The team tested the robot on real hardware, including the Ouster OS0 Rev7, Ouster OS1 Rev 7, and Livox MID-360. They ran the same test sequence on the new ROS 2 version and the old ROS 1 version. The result? The paths the robot drew were qualitatively identical. The robot navigated just as well as it did before, proving that the fixes didn't change how the robot thinks, they just stopped the new operating system from breaking it.
In short, moving a complex robot to a new system isn't just about translation; it's about understanding the new rules of the road. By fixing the delivery contracts and cleaning the buckets, the team saved the robot from a silent crash and got it back to exploring the world.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.