A Stitch in Time Saves Nine: Preserving Policy Compatibility Under Perception Updates in End-to-End Autonomous Driving
This paper proposes a lightweight model stitching approach that aligns latent representations between updated perception modules and frozen downstream policies, effectively preserving driving performance across diverse perception changes while significantly reducing adaptation time compared to traditional retraining methods.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Problem: Changing the Camera, Breaking the Brain
Imagine you have built a highly skilled self-driving car. This car has two main parts:
- The Eyes (Perception): A camera system that looks at the road and turns the view into a mental map (a "latent representation").
- The Brain (Policy): The decision-making software that looks at that mental map and says, "Turn left," "Stop," or "Speed up."
In modern "end-to-end" self-driving systems, the Eyes and the Brain are tightly married. They learn to speak the exact same language.
The Problem: What happens if you upgrade the camera? Maybe you add a new lens, change the sensor type, or retrain the camera software to be better at spotting pedestrians.
- The Result: The camera starts sending the Brain a message in a slightly different dialect. Even though the camera is seeing the same road, the "mental map" it sends looks different. The Brain, which was trained on the old dialect, gets confused. It might think a stop sign is a tree, or it might freeze up.
- The Old Solution: To fix this, engineers usually have to retrain the entire Brain from scratch. This is like hiring a new teacher to re-teach a student a new language. It takes weeks, requires massive amounts of data, and is incredibly expensive.
The Paper's Idea: The "Translator" (Model Stitching)
This paper proposes a much cheaper, faster solution called Model Stitching.
Instead of retraining the Brain, they ask: Can we just build a tiny, lightweight translator between the new Eyes and the old Brain?
Think of it like this:
- The Old Way: The new camera speaks "French." The old Brain only speaks "English." You fire the Brain and hire a new one that speaks French. (Expensive, slow).
- The New Way: You keep the English-speaking Brain. You install a small, cheap "translator" (the stitcher) between the camera and the brain. The translator instantly converts the French messages into English so the Brain doesn't have to change a thing.
How They Tested It
The researchers tested this "translator" under many different scenarios where the camera changed:
- Random Tweaks: Just changing the random starting numbers of the camera software.
- Different Architectures: Changing the camera from a "VAE" (a specific type of AI) to an "AE" (a different type).
- Different Sensors: Switching from using 4 cameras to 6 cameras, or using only LiDAR (lasers) instead of cameras.
- Different Worlds: Training the camera on real-world data (from the nuScenes dataset) but testing the car in a video game simulator (CARLA). This is the hardest test because the "worlds" look very different.
The Results: A Miracle of Efficiency
The paper found that this "translator" approach works incredibly well.
- Performance: In the hardest test (switching from real-world data to a simulator), the translator saved the car's performance. Without it, the car's driving score dropped to 34. With the translator, the score jumped back up to 89. This is nearly as good as retraining the whole system from scratch.
- Speed: This is the biggest win.
- Retraining the Brain: Takes 22.18 hours of computing time.
- Stitching (The Translator): Takes only 0.91 hours (less than an hour).
- Analogy: It's like the difference between rebuilding a house brick-by-brick versus just painting the front door to match the new neighbors.
- Data: Retraining requires the car to drive around in the simulator 70,000 times to learn. The translator needs zero extra driving. It just looks at a small batch of data once and figures out the conversion.
The "Secret Sauce": Why It Works
The paper suggests that even though the camera software changes, the important information (like "there is a car ahead" or "the road curves left") stays in a similar shape deep inside the AI's layers.
They call this the "Policy-Compatible Representation Stitchability Hypothesis."
- Simple Translation: If the change is small (like a different camera model), a simple Linear Stitcher (a basic math formula) works like a charm.
- Complex Translation: If the change is big (like switching from real life to a video game), they use a Convolutional Stitcher (a slightly more complex, flexible translator). This one is smart enough to handle the messy differences between the two worlds.
The Bottom Line
This paper proves that you don't need to throw away your self-driving car's "Brain" every time you upgrade the "Eyes." You can just sew on a tiny, cheap adapter that translates the new signals for the old brain.
This saves massive amounts of time, money, and computing power, making it much easier to keep self-driving cars safe and up-to-date as technology evolves. The authors plan to release their code so others can use this "translator" trick too.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.