VTO: Visual Tool Orchestration for Video Anomaly Detection
The paper proposes VTO, a process-supervised reinforcement learning framework that leverages a foundation model-driven cognitive evaluator and fine-grained step-wise supervision to enable multimodal agents to dynamically orchestrate specialized visual tools for improved video anomaly detection, validated by a new benchmark and toolset called VAD-Tool.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a detective trying to solve a mystery in a busy, chaotic city. You have a notebook and a pen, but the city is full of moving cars, falling objects, and people fighting. To solve the case, you can't just look at the whole scene and guess; you need to use specific tools. Maybe you need a magnifying glass to track a specific person, a speedometer to check if a car is speeding, or a fire detector to spot smoke. In the world of computers, this is called Video Anomaly Detection. It's the job of teaching machines to spot weird or dangerous things in video footage, like a fight breaking out or someone falling down.
For a long time, computers tried to do this by memorizing patterns, like a student cramming for a test. They would look at a video and say, "This looks like a fight!" But if the fight happened in a new place or looked slightly different, the computer would get confused and fail. Recently, scientists tried giving computers a "brain" (a Large Language Model) that could talk and think. They hoped the computer could figure out which tool to use and when. But here's the problem: these smart computers often stop thinking and give their answer immediately after seeing one small problem. They miss the bigger picture, like how a small fight could turn into a stampede. They need a way to learn how to keep investigating, step-by-step, without giving up too soon.
This is where the new paper comes in. The researchers, led by Rui Wang and Mengshi Qi from Beijing University of Posts and Telecommunications, built a new system called VTO (Visual Tool Orchestration). Think of VTO as a strict but brilliant coach for a detective robot. Instead of just letting the robot guess, the coach watches every single step the robot takes. If the robot picks the wrong tool, or stops investigating before it's sure, the coach gives it a "thumbs down" right then and there. If the robot follows a perfect logical chain, checking one clue after another, the coach gives it a "thumbs up."
The team created a special training ground called VAD-Tool, which is like a giant playground with 12 different "magic tools" for the robot to learn. These tools can do things like count crowds, recognize faces, spot weapons, or detect fire. The researchers taught their robot using a method called "process-supervised reinforcement learning." In simple terms, this means they didn't just grade the robot on the final answer; they graded the process. They made sure the robot learned that stopping too early is a mistake. They even used a super-smart AI (a 72-billion-parameter model) to act as a judge, checking if the robot's reasoning made sense at every step.
The results were impressive. When they tested their new VTO system, it got much better at solving these complex, multi-step mysteries than the old methods. While other systems often gave up after finding one clue, VTO kept going, connecting the dots until it found the whole story. In their tests, the new system improved the accuracy of picking the right tools and the right order by up to 10.2% compared to the previous best methods. It even beat some of the biggest, most powerful AI models out there, proving that teaching a computer how to think step-by-step is more important than just making it bigger. The paper suggests that by forcing the AI to follow a strict, logical path and rewarding it for not giving up, we can build safer, smarter systems to watch over our cities.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.