Resource-Aware Small-UAV Perception under Annotation and Compute Constraints
This paper proposes and evaluates a resource-aware perception strategy for small UAV detection that combines confidence-guided model routing and difficulty-ranked replay to effectively optimize performance under simultaneous constraints on inference-time computation and annotation budgets.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the quiet spaces between the ground and the clouds, a new kind of surveillance is taking shape. Small, unmanned aircraft are becoming common tools for everything from inspecting power lines to monitoring wildlife, but they are also a challenge for safety and security. To keep an eye on them, we need cameras that can spot a tiny, distant speck against a vast, shifting sky. This is a difficult task for a computer. A drone far away might occupy only a few pixels on a screen, easily lost in the visual noise of clouds, buildings, or tree branches. Furthermore, the computers that run these cameras are often small and battery-powered, meaning they cannot afford to run heavy, complex software on every single image they capture. At the same time, teaching these computers to recognize drones usually requires thousands of labeled photos, a resource that is often scarce when a new monitoring mission begins. The core question is how to build a system that sees clearly without needing too much power or too many examples to learn from.
Researchers have tackled this problem by treating the limits of power and data not as obstacles to be overcome, but as conditions to be managed. They tested a strategy on a standard set of drone images, using a popular type of detection software that comes in different sizes, from very light and fast to heavier and more detailed. Their first discovery was straightforward: the size of the image matters more than the size of the software. When they fed the computer images at a higher resolution, allowing it to see more detail, the system became significantly better at finding the drones, even when using a smaller, faster program. The best single setup they found used a medium-sized program running on high-resolution images, which caught nearly 91 percent of the drones in their test set. However, running a heavy program on every single frame is too slow for a small, battery-powered device.
To solve the speed problem, the team introduced a two-step process that acts like a filter. First, a very lightweight, fast program scans every image. If it is confident it sees nothing or sees something clearly, it moves on. But if the fast program is unsure, it passes that specific image to a larger, more powerful program for a second look. This approach, which they tested by routing about 36 percent of the images to the stronger program, improved the system's ability to find drones without slowing it down to the speed of the heavy program alone. The system became much better at catching the drones it had previously missed, raising its overall success rate, while still running at a speed of about 16 frames per second. Crucially, the researchers proved that this improvement came from the smart choice of which images to re-examine, not just from looking at more images. When they randomly sent images to the stronger program instead of using the confidence-based filter, the system performed worse.
The study also addressed the problem of limited training data. In many real-world scenarios, a new monitoring task might start with only half the photos needed for a full training session. The researchers tested a method to make the most of these scarce labels by focusing the computer's learning on the images it found most difficult. They ranked the training photos based on how many drones the computer missed, how poorly it located them, or how small the drones were. By replaying these difficult examples more often during training, the system learned slightly better than if it had just repeated all the photos equally. This benefit was most noticeable when the system was trying to find the tiniest drones, improving its ability to spot them by a small but meaningful margin. However, this trick worked best when the computer already had a decent foundation of knowledge; when the training data was extremely scarce, the system was too confused to know which examples were truly difficult.
The researchers also checked how well their system held up against bad weather and image distortions, such as blur or noise. The system remained remarkably stable, retaining over 96 percent of its accuracy even when the images were corrupted. They even tested whether a large, advanced artificial intelligence model could help by looking at the whole picture to find drones, but found that these models were unreliable for this specific task without the initial help of the camera-based detector. The final picture that emerges is one of balance. The most effective approach for spotting small drones on limited hardware is not to use the biggest possible model, but to use a smart combination of a fast first pass and a careful second look, supported by high-resolution images and a training process that focuses on the hardest examples. This resource-aware strategy allows a small computer to see what it needs to see, proving that careful management of limited resources can be just as powerful as raw computing power.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.