A Cost-Aware, Paired Protocol for Auditing Dynamic Tool Synthesis in Agentic Video Question Answering
This paper introduces a cost-aware, paired auditing protocol that jointly evaluates accuracy and inference cost to reveal that the Dynamic-SAGE framework, which synthesizes composite tools for agentic VideoQA, significantly improves accuracy and reduces reasoning turns while increasing token usage and overall cost, demonstrating the necessity of multi-axis evaluation over scalar metrics.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to solve a complex mystery by watching a long, complicated video. You have a toolbox full of basic tools: a magnifying glass, a microphone, a search engine, and a notepad.
The Old Way (Static-SAGE)
In the traditional approach, every time you get a new question about the video, you have to start from scratch. Even if you need to do the same three steps over and over again (like "find the scene," "listen to the audio," and "read the transcript"), you have to manually pick up each tool, use it, put it down, and then pick up the next one for every single question. It's like having to walk to the kitchen, get a knife, chop an onion, walk back, and then do it again for every single sandwich you make. It works, but it's slow and repetitive.
The New Idea (Dynamic-SAGE)
The researchers behind this paper asked: "What if we could build a custom, multi-step machine that does those three steps automatically?"
They created a system called Dynamic-SAGE. Before the system starts answering questions, it watches a set of practice videos. It notices patterns, like "Oh, every time someone asks about a specific event, the system always listens to the audio and then looks at the frames." So, it builds a new, custom "super-tool" that combines those steps into one button.
Now, when the system faces a new question, instead of clicking three separate buttons, it just hits one "super-button" that does the whole job instantly.
The Catch: The "Cost-Aware" Audit
Here is the tricky part. The researchers realized that just checking if the answer is right isn't enough. You also need to know if the system is working harder or smarter.
They invented a new way to test this, like a financial audit for a factory. They didn't just ask, "Did the new machine make better sandwiches?" They asked:
- Did the sandwiches taste better? (Accuracy)
- Did the machine use fewer steps? (Efficiency)
- Did it use more electricity or expensive ingredients? (Cost)
What They Found
When they ran the audit, the results were a mix of good news and a twist:
- The Good News: The new system got the answers right 7.5% more often than the old one. It also stopped "wandering around" so much, reducing the number of times it had to ask for help by about 28%. It was like the chef stopped walking back and forth to the kitchen.
- The Twist: Even though the chef took fewer steps, the "electricity bill" (the cost of the computer processing) went up by 26%.
- Why? Because the "super-tools" are heavy. When the system hits that one "super-button," it actually runs a lot of complex calculations inside that single button. It's like switching from a bicycle to a sports car: you get to your destination faster and with fewer gear shifts, but you burn more gas per mile.
Where It Works Best
The new system is a superhero for visual questions (like "What color was the car?") and open-ended questions (like "Why did the character leave?"). It helps the most when the video is long and complicated.
However, it doesn't help much with questions that rely purely on speech or a mix of sound and sight. In those cases, the new tools didn't make a big difference.
The Bottom Line
The paper concludes that building these custom "super-tools" is a smart move because it makes the system smarter and more accurate, even if it costs a bit more to run. It's a trade-off: you pay a little more in "gas money" (computing cost) to get a much better driver (the AI) who makes fewer mistakes.
The researchers warn that this system still struggles with the hardest, most confusing questions, but for the rest, it's a clear upgrade that saves time and effort, even if the price tag is slightly higher.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.