ARMOR++: Agentic Orchestration of a Multi-Domain Primitive Set for Transferable Attacks on Deepfake Detectors
This paper introduces ARMOR++, a robust multi-agent framework that leverages Qwen2.5-VL and Qwen3 to orchestrate a diverse set of adversarial primitives, achieving significantly higher transferability and attack success rates against deepfake detectors compared to existing state-of-the-art methods under strict black-box constraints.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the internet as a giant, bustling art gallery where anyone can paint a picture. For a long time, if you wanted to know if a painting was real or a forgery, you just looked at it with your eyes. But recently, a new kind of magic brush appeared—artificial intelligence—that can paint faces so perfectly that even experts struggle to tell the difference. This has created a high-stakes game of cat and mouse. On one side, we have "detectives" (AI models) trained to spot the tiny, invisible flaws left behind by these magic brushes. On the other side, we have "hackers" trying to trick the detectives by adding just enough noise to a fake face to make it look real to the detective, without actually changing how the face looks to a human. This isn't just about art; it's about trust. If we can't tell what's real, we can't trust our news, our security cameras, or even our own eyes. The big question is: how do we know if our digital detectives are actually reliable, or if they are just easily fooled by a clever trick?
Enter ARMOR++, a new, super-smart "hacker" team designed to test these detectives. Think of deepfake detectors like security guards at a club. Some guards only look at the shape of a person's face (like a Convolutional Neural Network), while others look at the whole picture and how the pieces fit together (like a Vision Transformer). The problem is that a trick that works on one type of guard often fails on the other. Previous attempts to trick these guards were like throwing a single type of rock at a wall; sometimes it worked, but often the wall was too strong or made of a different material.
ARMOR++ is different because it doesn't just throw one rock; it sends in a whole team of specialists who talk to each other. The team is led by two very smart "agents" (computer programs based on large language models). The first agent, the Analyst, looks at the fake face and uses its "eyes" to find the weird spots—like a blurry edge or a strange texture—and says, "Hey, attack here!" The second agent, the Conductor, acts like a coach. It doesn't just pick one attack; it manages a squad of five different "attackers," each with a unique superpower:
- The Blender: Adds a smooth, dense layer of noise everywhere.
- The Pincher: Tweaks just a few specific pixels to confuse the guard.
- The Shifter: Gently warps the face, like stretching a rubber sheet.
- The Frequency Wizard: Changes the "colors" of the image in a way humans can't see but computers can (like turning a radio dial).
- The Block Shuffler: Cuts the image into puzzle pieces and swaps them around.
What makes ARMOR++ special is that it plays a game of "guess and check" without ever asking the target guard for help. It practices on a set of practice guards (surrogate models) that it knows inside out. It tries out different combinations of its five attackers, listens to its smart agents, and adjusts its strategy in real-time. If an attack isn't working, the team doesn't give up; they change the rules, try a different mix of attacks, and keep going until they find the perfect combination that fools the target guard.
The paper finds that this team-based approach is incredibly effective. When tested on a massive collection of fake faces (the AADD-2025 benchmark), ARMOR++ managed to fool the most advanced "blind" guards (the ones it had never seen before) about 44.3% of the time on low-quality images and 32.1% of the time on high-quality images. To put that in perspective, the previous best team (called ARMOR) only fooled them about 39.6% and 28.3% of the time, and standard "single-attack" methods barely broke 20%.
The researchers also checked if these tricks still worked if the guards had some basic training to resist attacks (like "adversarial training"). Even then, ARMOR++ was the most successful, still fooling the guards about 19.8% of the time, while other methods dropped to near zero. This suggests that even our best current security systems have a significant "blind spot" when it comes to these smart, multi-pronged tricks. The authors conclude that while our detectors are getting better at spotting fakes, they are still not reliable enough to be trusted completely, especially when facing an attacker that can think, adapt, and use many different types of tricks at once. It's a wake-up call: in the race between fake creators and real detectors, the detectors still have a lot of work to do.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.