Deadline-Aware Hardening of Real-Time Object Detection Against Candidate-Inflation Latency Attacks
This paper proposes a retraining-free, deployment-selectable mechanism that caps the number of candidates entering non-maximum suppression to a deadline-calibrated bound, thereby mitigating candidate-inflation latency attacks and ensuring real-time deadline integrity across diverse hardware and detector architectures while revealing that bounding suppression alone is necessary but insufficient due to significant decoding overhead.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the world of autonomous vehicles and security cameras, seeing is not enough; seeing in time is everything. A computer vision system designed to spot pedestrians or traffic signs must do more than simply identify them correctly. It must deliver that identification before the next moment arrives. If a car traveling at highway speeds receives a warning about a hazard a fraction of a second too late, the result is not merely a slower response, but a potential catastrophe. This requirement creates a strict deadline for every single image the system processes. If the computer takes too long to finish its work on one picture, the pipeline falls behind, and the output becomes stale, describing a scene that has already passed.
For years, researchers have focused on making these systems faster and more accurate, often measuring success by average speed. However, in a real-time system, the average can be misleading. A system might be incredibly fast most of the time but occasionally freeze for a long duration. In a safety-critical application, that single slow moment is a failure. Furthermore, these systems are not just vulnerable to random glitches; they can be targeted by attackers who do not try to trick the computer into seeing the wrong object, but rather into working so hard that it runs out of time. This paper explores a specific type of attack where an adversary subtly alters an image to force the computer to generate an overwhelming number of potential detections, causing it to miss its deadline. The researchers then propose a simple, practical way to stop this without needing to retrain the computer's brain.
The core of the problem lies in how these detectors work. When a camera captures an image, the software scans it and produces a massive list of potential objects, each with a confidence score. To turn this chaotic list into a clean set of final detections, the system uses a process called non-maximum suppression. Imagine a crowded room where many people are shouting out the same name; this process filters out the duplicates and keeps only the loudest, most confident voices. Under normal conditions, this filtering is quick. However, an attacker can craft an image that tricks the system into generating tens of thousands of potential objects instead of a few dozen. The filtering process then has to compare every single one of these thousands of candidates against every other one. This creates a computational explosion. The more candidates the attacker forces the system to consider, the longer the filtering takes, eventually causing the system to miss its deadline and fail to deliver a result in time.
The researchers tested this threat on a real-time object detection system running on powerful hardware, specifically designed to handle video streams at thirty frames per second. They found that a standard, unmodified system could be easily overwhelmed. When they fed the system images designed to trigger this overload, the time it took to filter the candidates jumped from a fraction of a millisecond to hundreds of milliseconds. Even on the fastest hardware they tested, the system failed to meet the deadline for the filtering stage if the number of candidates was left unchecked. However, the study confirmed that simply using faster hardware or a more efficient software version of the filtering process was not enough to solve the problem on its own. While these improvements made the system faster, they did not stop the attacker from controlling the workload. The attacker could still force the system to do so much work that even the fastest machine would stumble if no limit was placed on the input.
To solve this, the researchers introduced a strict limit on the number of candidates allowed to enter the filtering stage. Instead of letting the system process every single potential object the image generated, they capped the number at a specific, manageable level. If the system produced more candidates than this limit, it simply selected the most promising ones and discarded the rest before the heavy filtering began. This approach acts as a safety valve, ensuring that the amount of work the system must perform never exceeds a known, safe maximum. The researchers carefully measured the cost of this safety measure. They found that by limiting the candidates to one thousand and twenty-four, the system could handle the filtering stage well within the deadline for that specific stage, reducing the latency to just 4.03 ms. However, the study revealed a critical nuance: even with this cap in place, the defended requests still missed the overall end-to-end deadline. This was not solely due to the attack, but because other bottlenecks, such as the time required to decode the image itself, consumed the remaining time budget. In fact, the researchers found that clean images without any attack also missed the overall deadline 92.7% of the time when using lossless formats, indicating that the decoding process was a major bottleneck regardless of the attack. The trade-off for this protection was an almost imperceptible drop in accuracy, measured at a tiny fraction of a percent, which is negligible for practical use.
The study went further to ensure this solution was robust across different scenarios. They tested the method on two different types of camera sensors and with different software backends, including those running on standard computers and those running on smaller, energy-efficient edge devices. In every case, the limit held firm for the filtering stage, preventing the attacker from inflating the workload beyond the cap. Even when the hardware was under stress from heat or when the system was running on a less powerful board, the capped approach prevented the filtering stage from stalling. However, the researchers emphasized that while the cap successfully controlled the filtering stage, it did not guarantee that the entire pipeline would meet the deadline. They found that once the filtering was controlled, the next bottleneck was often the time it took to decode the image itself. This means that while limiting the candidates is a necessary step to protect the system from this specific attack, it is not a complete cure-all; the entire pipeline must be monitored to ensure the deadline is met.
The authors argue that this method of setting a hard limit on the workload is a crucial step for deploying real-time vision systems in the real world. It shifts the control of the workload from the attacker back to the system administrator. By defining a maximum number of candidates based on the system's speed and the required deadline, a deployment can guarantee that it will never be forced to do more work than it can handle during the filtering stage. The paper concludes that while faster hardware and better algorithms are helpful, they are not sufficient on their own. A real-time system needs a hard boundary on the work it is asked to do. Without such a boundary, an attacker can always find a way to overwhelm the system. With it, the system remains reliable in its processing of the filtering stage, delivering its results on time for that specific component, even when the world around it tries to break it, though the overall system deadline depends on managing all other stages as well.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.