GuideMe: Multi-Domain Task Guidance and Intervention in Streaming Video
This paper introduces GuideMe, the first multi-domain benchmark for streaming video designed to evaluate and train multimodal Large Language Models on closed-loop interactive task guidance, revealing that while current models excel at providing instructions, they significantly struggle with real-time error detection and corrective feedback.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to learn how to assemble a complex piece of furniture or cook a complicated meal. You have a very smart, knowledgeable friend (an AI) who has read every manual and watched every cooking show.
The Problem:
Right now, if you ask this friend for help, they are great at giving you a list of instructions after you've finished the whole thing. They can tell you, "Well, you did step 1, then step 2, and then you made a mistake." But they are terrible at being a real-time coach. They can't watch you while you are working, spot that you are about to use the wrong screwdriver, and immediately say, "Wait! Stop! That's the wrong tool!"
Most current AI models are like a movie critic who only writes a review after the movie ends. They aren't good at being a director who yells "Cut!" while the scene is being filmed.
The Solution: "GuideMe"
The authors of this paper built a new testing ground called GuideMe. Think of it as a giant gym for AI coaches.
- The Workout: They created a massive library of over 2,400 videos (like cooking, fixing things, daily tasks, and fitness) totaling over 223 hours.
- The Drill: They didn't just label the videos; they turned them into interactive conversations. In this gym, the AI has to:
- Give the next instruction.
- Watch the user do it.
- Crucially: If the user makes a mistake, the AI must spot it instantly and say, "No, that's wrong, try this instead."
- If the user is doing fine, the AI must know to stay quiet and not interrupt.
The Big Discovery
The researchers put many different AI models (both famous commercial ones and open-source ones) into this gym to see how they performed. The results were surprising and a bit disappointing:
- The "Talkers": Some AIs were too eager. They kept shouting instructions even when the user was doing fine, or they gave advice at the wrong time. It was like a coach who never stops talking, even when you're just stretching.
- The "Silencers": Other AIs were too shy. They watched the user make huge mistakes but said nothing, waiting for the user to ask for help. They were like a coach who just sits on the bench and watches you fail.
- The Gap: The AIs were actually quite good at saying, "Here is what you should do next." But they were terrible at the hard part: watching what you are doing, realizing it's wrong, and fixing it on the fly.
How They Measured Success
To grade the AI, they used a three-part scorecard:
- Timing: Did the AI speak at the exact right moment, or was it too early/late?
- Behavior: Did the AI know when to stay silent and when to speak? (This was the hardest part).
- Quality: When the AI did speak, was the advice actually helpful and correct?
The Bottom Line
The paper concludes that while AI is getting very good at describing videos and giving step-by-step lists, it is still far from being a reliable, real-time coach. It can tell you the recipe, but it can't yet watch you cook and stop you from burning the toast. The authors hope this new "gym" (GuideMe) will help future AI developers train their models to become true interactive coaches rather than just passive narrators.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.