DETOUR: A Practical Backdoor Attack against Object Detection
The paper proposes DETOUR, a practical backdoor attack against Object Detection Transformers that leverages a "trigger radiating effect" by training models with semantic triggers rescaled to various sizes and extracted from real-world objects under diverse fields of view, thereby ensuring robust activation across different spatial configurations and viewpoints.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a very smart security camera system (an "Object Detection" model) that is supposed to spot cars, people, and animals. Its job is to look at a picture, find everything in it, and say, "That's a car," or "That's a dog."
This paper introduces a new way to trick this camera system, which the authors call DETOUR.
Here is the breakdown of how it works, using simple analogies:
1. The Problem with Old Tricks
Previous attempts to hack these cameras relied on "invisible" tricks. Imagine trying to trick the camera by adding a tiny, almost invisible speck of dust to a photo.
- The Flaw: In the real world, cameras aren't perfect. If you print that tiny speck on a piece of paper and hold it up to a camera, the camera might miss it because of lighting, distance, or the angle. It's like trying to whisper a secret to someone wearing noise-canceling headphones; the message gets lost.
- The Limitation: These old tricks also only worked if the "speck" was in the exact same spot every time. If the camera moved slightly, the hack stopped working.
2. The New Strategy: The "Magic Mug"
The authors of this paper decided to stop trying to be invisible. Instead, they used something obvious and real: a coffee mug.
- The Trigger: They use a picture of a real-world object (like a mug) as the "trigger." It's not a weird, invisible pattern; it's a normal object you might see on a table.
- Why it works: Because it's a real object, cameras can easily see it, no matter how the lighting changes or how far away it is. It's like shouting a secret instead of whispering it; the camera definitely hears it.
3. The "Radiating Effect" (The Secret Sauce)
The most interesting discovery the authors made is something they call the Trigger Radiating Effect (TRE).
- The Analogy: Imagine you drop a stone in a pond. The ripples don't just stay where the stone hit; they spread out in all directions.
- How it applies: The authors found that if they trained the camera to recognize the "Mug" trigger, the camera didn't just react when the mug was in the exact center of the photo. It reacted even if the mug was slightly to the left, right, up, or down. The "hack" radiated out from the trigger location.
- The Upgrade: By training the camera with the mug in many different spots and at many different sizes (zoomed in, zoomed out), the "ripples" cover the whole pond. Now, no matter where the mug appears in the photo, the camera gets tricked.
4. What Happens When the Hack Works?
Depending on what the attacker wants, the camera starts doing one of three things when it sees the mug:
- Misclassification: It sees a dog, but because the mug is there, it screams, "That's a person!"
- Disappearance: It sees a car, but because the mug is there, it pretends the car isn't there at all.
- Hallucination: It sees an empty street, but because the mug is there, it invents a fake car or person out of thin air.
5. Why is this a Big Deal?
The authors tested this against six different "hacking goals" (like making the camera miss things or see things that aren't there).
- The Result: Their method (DETOUR) was 36% more effective at tricking the camera than previous methods, and 70% better at working no matter where the trigger was placed.
- The Safety: Crucially, when the mug isn't in the picture, the camera still works perfectly. It doesn't break the camera; it just adds a hidden "backdoor" that only opens when the mug is present.
Summary
Think of DETOUR as teaching a security guard to ignore everything in a room unless they see a specific red mug.
- Old hacks tried to teach the guard to ignore things if they saw a tiny, invisible dot (which the guard often missed).
- DETOUR teaches the guard to react to a big, obvious red mug.
- Even better, the guard learned that if the mug is anywhere in the room (not just on the desk), the rule applies.
- And if there is no mug? The guard does their job perfectly, spotting all the real cars and people.
The paper proves that by using real-world objects and understanding how the camera's "brain" spreads its attention, attackers can create a much more reliable and practical way to fool these systems.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.