← Latest papers
💻 computer science

AdvSerial: Physical Adversarial Attacks on Infrastructure-mounted Pedestrian Detectors via Semantic Feature Suppression

The paper introduces AdvSerial, a dynamic 2D-3D joint optimization framework that generates continuous high-angle physical adversarial patches to effectively suppress pedestrian semantic features and achieve high attack success rates against infrastructure-mounted detectors while evading existing defenses.

Original authors: Yuanhao Huang, Yilong Ren, Jinlei Wang, Xuesong Bai, Jinchuan Zhang, Haiyang Yu

Published 2026-07-21
📖 7 min read🧠 Deep dive

Original authors: Yuanhao Huang, Yilong Ren, Jinlei Wang, Xuesong Bai, Jinchuan Zhang, Haiyang Yu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a world where your eyes and brain are replaced by a super-smart camera computer. This is the reality of modern "AI vision," a technology that helps self-driving cars see pedestrians, helps security cameras count crowds, and keeps an eye on busy highways. These systems are incredibly good at spotting people, but they have a secret weakness: they can be tricked. Just like a human can be fooled by a clever optical illusion, these AI cameras can be confused by special patterns designed to look like noise to us but act like a "blindfold" to the computer. This field of study is called "adversarial attacks." It's not about breaking the camera with a hammer; it's about painting a picture that makes the AI's brain short-circuit. Why does this matter? Because if a security camera at a busy train station or a traffic light can be tricked into not seeing a person, the whole system fails, potentially leading to dangerous accidents or security breaches.

Now, meet AdvSerial, a new "magic cloak" designed by researchers to test just how vulnerable these AI cameras really are. While previous attempts to trick AI focused on flat, street-level views (like a camera at eye level), this new method targets the cameras mounted high up on poles and buildings—the kind used in smart cities and on highways. These high-angle cameras see people from above, which distorts their shape and makes them harder to trick. The researchers built a digital "sewing machine" that takes a special, chaotic pattern and stitches it onto 3D models of clothes. They then taught this pattern to survive the weird angles and movements of real life. When they tested it, the results were startling: in a physical world experiment, the cloak successfully hid people from a popular AI detector called YOLO-v5 about 74.8% of the time. Even more impressively, it reduced the camera's confidence in seeing a person from a solid 84.30% down to a shaky 39.38%. The best part? This trick worked even against advanced defenses that try to spot "weird" patterns or look at video over time to catch the trick.

The Story of the "Invisible" Jacket

Let's dive into how this works. Imagine you are trying to hide a person from a security guard who is looking down from a tall tower. If you just put a big, ugly sticker on their shirt, the guard might spot the sticker and ignore it. If you try to paint a pattern on their shirt that looks like a normal shirt but confuses the guard, the pattern might get stretched and twisted when the person walks, making it look like a mess and giving the game away.

The researchers, led by Yuanhao Huang and Yilong Ren, realized that to truly hide a person from these high-up cameras, you couldn't just make a flat sticker. You needed a dynamic 3D suit. They created a system called AdvSerial that does three main things:

  1. The Digital Sewing Machine (Feature Smooth Quilting): Imagine you are making a quilt out of many small squares of fabric, each with a crazy, confusing pattern. If you just sew them together, the seams (the lines where the squares meet) will be obvious. A smart camera might see those sharp lines and say, "Aha! That's a fake pattern!" To fix this, the researchers invented a clever way to sew the squares. Instead of just matching the colors, they looked at where the camera's brain was most sensitive. They stitched the squares together in the "blind spots" of the camera's brain—places where the camera doesn't pay much attention. This is called Feature Smooth Quilting. It makes the seams invisible to the AI, even though they are there, ensuring the pattern looks like one smooth, confusing surface.

  2. The Time-Traveling Mannequin (Serial Frame Loss): Most previous tricks worked for a single photo. But in the real world, people move! They walk, run, and swing their arms. If the pattern only works when you stand still, it's useless for a security camera watching a video. The researchers trained their pattern on a 3D mannequin that was constantly moving, changing poses, and spinning around. They used a special math rule called Serial Frame Loss to make sure the pattern kept working throughout the whole video. It's like teaching a magician's trick not just to work for one second, but to keep working for a whole minute, no matter how the magician moves. This prevented the "flickering" effect where the AI sees the person for a second, then loses them, then sees them again.

  3. The High-Angle Test: They didn't just test this on a computer screen. They printed the patterns onto real clothes and had people walk around in front of real security cameras mounted 3 to 5 meters high. They tested this in different weather (sunny, rainy, foggy) and at different distances (from 5 to 30 meters away).

The Results: A Masterclass in Hiding

The results were a wake-up call for the security world. When they tested their "magic cloak" against eight different types of AI detectors (including the popular YOLO series and others), the results were consistent and scary:

  • The Big Win: Against the YOLO-v2 detector, the cloak worked 89.71% of the time. Against the more modern YOLO-v5 in the real world, it still worked 74.8% of the time.
  • Confidence Crash: Before the attack, the cameras were very sure they saw a person (84.30% confidence). After the attack, that confidence plummeted to 39.38%. The camera was basically saying, "I'm not sure if that's a person or just a shadow."
  • Beating the Defenses: The researchers tried to break their own trick using two advanced defenses. One defense (NapGuard) looks for sharp, unnatural lines in patterns. The other (Sparse4D-v3) looks at video over time to see if things are moving strangely. The cloak beat both. It was so smooth that the line-detecting defense missed it, and it was so consistent over time that the video-checking defense couldn't find a glitch.

Why This Matters (And What It Doesn't)

This paper doesn't say that AI is broken forever. Instead, it acts like a stress test. It shows that the cameras we rely on for safety in our cities have a specific blind spot: they struggle to see people wearing these specific, carefully designed patterns when viewed from high angles.

The researchers found that the trick works best when the person is standing or moving slowly. When people run or swing their arms wildly, the pattern gets stretched, and the camera might catch a glimpse of the person again (the success rate dropped a bit, but it was still much lower than normal). This tells us that while the attack is powerful, it's not a magic "invisibility cloak" that works 100% of the time in every situation.

The most important takeaway is that the old ways of testing security (just taking a photo and seeing if the AI gets confused) aren't enough. We need to test how these systems handle moving people, changing angles, and real-world weather. The researchers suggest that to fix this, future security systems need to be smarter about how they track movement and how they handle 3D shapes, not just 2D pictures.

In short, AdvSerial is a powerful tool that exposes a hidden weakness in our smart city cameras. It proves that if you know how the AI's brain works, you can design a pattern that makes it "go blind" to people, even from high above. It's a reminder that as we build smarter cities, we also need to build smarter defenses to keep them safe.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →