CompAgent: An Agentic Framework for Visual Compliance Verification
This paper introduces CompAgent, the first agentic framework that enhances Multimodal Large Language Models with dynamic tool selection and structured reasoning to achieve state-of-the-art performance in visual compliance verification across complex policy domains.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are the editor-in-chief of a massive, 24-hour news channel that broadcasts millions of images every single day. Your job is to make sure nothing offensive, dangerous, or against the rules gets on air.
In the past, you had two main ways to do this:
- The Specialist: You hired a team of experts who only looked for one thing (like a "Nudity Detective" or a "Violence Spotter"). They were fast but couldn't handle complex situations. If a picture showed a soldier with a gun, the Violence Spotter might scream "STOP!" without realizing it's a news photo about a war, not a threat.
- The Generalist: You hired one incredibly smart, well-read editor (an AI) who knew everything about the world. But this editor was sometimes too literal. They might see a knife in a cooking show and think it's a crime, or miss a subtle joke that was actually harmful.
CompAgent is a new, revolutionary way to solve this problem. Think of it not as a single person, but as a high-tech detective agency working together.
The Three-Part Detective Team
CompAgent breaks the job down into three distinct roles, working like a well-oiled machine:
1. The Case Manager (The Planning Agent)
Imagine a seasoned detective who doesn't look at the crime scene photos themselves. Instead, they read the Rule Book (the compliance policy) and ask: "What kind of evidence do we need to solve this case?"
- If the rule says "No weapons," the Case Manager calls the Ballistics Expert.
- If the rule says "No hate speech," they call the Linguist.
- If the rule is about "Self-harm," they call the Medical Analyst.
The Case Manager is smart because they don't waste time. They don't ask the Ballistics Expert to look at a picture of a puppy. They only call the specific tools needed for that specific rule.
2. The Toolkit (The Tool Suite)
This is the team of specialists the Case Manager calls. They are like a Swiss Army knife of AI tools:
- The Eyes: They can spot objects (is that a gun or a toy?), faces (is that a child?), and text (what does the sign say?).
- The Context Readers: They can summarize the whole scene in a sentence.
- The Safety Checkers: They have specialized training to spot things like "disturbing content" or "illegal activities."
These tools are like the police, the forensics lab, and the legal team all rolled into one. They gather the hard facts.
3. The Judge (The Compliance Verification Agent)
Once the Case Manager gathers all the reports from the specialists, they hand the file to the Judge.
The Judge looks at the original picture and all the reports from the team. They are the only one who makes the final call.
- The Ballistics Expert says: "That's a gun."
- The Context Reader says: "It's a soldier in a parade."
- The Judge says: "Okay, a gun is usually bad, but in this specific context (a parade), it's allowed. Safe."
Or, in a tricky case:
- The Specialist says: "I see a person with cuts on their arm."
- The Context Reader says: "This is a poster for mental health awareness."
- The Judge says: "Even though the intention is good, the image of the cuts is so graphic it might hurt someone watching. Unsafe."
Why is this better than the old ways?
1. It's Flexible (Like a Chameleon)
Old systems were like a locked door; if the rules changed, you had to rebuild the whole door. CompAgent is like a chameleon. If the company changes the rules tomorrow (e.g., "Now we also need to check for AI-generated deepfakes"), the Case Manager just picks up a new tool from the toolbox. No need to retrain the whole team.
2. It Explains Its Work (The Paper Trail)
If a human editor rejects a photo, they write a note explaining why. CompAgent does the same. It doesn't just say "Unsafe." It says: "I flagged this as Unsafe because the Object Detector saw a knife, and the Text Detector read a threatening message, which violates Rule #4." This makes it easy for humans to trust the AI.
3. It's Smarter at Nuance
The paper tested CompAgent on thousands of images. It beat the "Specialists" and the "Generalist" AI.
- Example: A picture of a soldier with a helicopter.
- Old AI: "Weapon detected! Unsafe!" (Too simple).
- CompAgent: "Detects soldier + helicopter. Checks policy: 'Legal military operations are allowed.' Safe." (Understands the context).
The Bottom Line
CompAgent is like upgrading from a single security guard with a flashlight to a smart security system with a control room. The control room (the Agent) decides which cameras (tools) to turn on based on the specific threat, gathers the evidence, and makes a final, well-reasoned decision.
It's faster, cheaper (because it doesn't need to be retrained every time rules change), and much better at understanding the difference between a scary movie scene and a real crime. It's the future of keeping the internet safe.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.