ProcObject-10K: Benchmarking Object-Centric Procedural Understanding in Instructional Videos
This paper introduces ProcObject-10K, the first benchmark designed to evaluate object-centric reasoning and temporal grounding in instructional videos, revealing that current multimodal models rely heavily on linguistic priors rather than fine-grained object dynamics while demonstrating that object-centric fine-tuning can effectively bridge this gap.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are watching a cooking show. A standard AI might look at the video and say, "The chef is chopping onions, then frying them." It sees the actions (chopping, frying) but doesn't really understand what is happening to the ingredients. It doesn't realize that the onion changed from a solid, crunchy state to a soft, translucent one, or that if the chef had forgotten to add oil, the onions would burn instead of fry.
This paper, ProcObject-10K, argues that current AI is too focused on the "dance moves" (actions) and not enough on the "dancers" (objects) and how they change.
Here is a simple breakdown of what the researchers did:
1. The Problem: The "Blind" Chef
The authors say that existing AI benchmarks for instructional videos are like asking a student to describe a movie by only listing the plot points, without noticing how the characters' emotions or the setting changed.
- The Gap: AI models can often guess the right answer to a question like "What happens next?" because they have read millions of recipes on the internet (linguistic priors). They know onions usually get fried.
- The Failure: However, if you ask the AI to point to the exact moment in the video where the onion changed from raw to cooked, or to explain why a mistake happened (e.g., "The butter didn't melt because it was too cold"), the AI fails. It can't find the visual evidence. It's like a student who memorized the answer key but didn't read the textbook.
2. The Solution: A New "Gym" for AI (ProcObject-10K)
To fix this, the researchers built a new testing ground called ProcObject-10K. Think of this as a specialized gym designed to train AI to pay attention to objects and their state changes.
- The Dataset: They gathered over 1,700 video clips (cooking, assembly, lab work) and created 10,500 questions.
- The Questions: Instead of just asking "What did the chef do?", they ask:
- Precondition: "What had to happen to the egg before it could be whisked?" (It had to be cracked).
- State Evolution: "How did the egg change from the bowl to the microwave?" (It went from liquid to solid curd).
- Counterfactuals: "What would have happened if the garlic wasn't peeled?" (It would be chunky and bitter).
- Mistakes: "What went wrong with the butter?" (It stayed in chunks instead of melting).
- The Evidence: Crucially, every answer must be backed up by specific time stamps in the video. The AI must prove where it saw the change.
3. The Test Results: The "Answering-Grounding Gap"
The researchers tested 13 of the smartest AI models available (like GPT-4, Gemini, and others) on this new gym.
- The Result: The models were great at writing plausible answers (scoring well on language), but terrible at pointing to the video evidence.
- The Metaphor: Imagine a detective who can write a perfect story about who committed the crime based on the news, but when asked to point to the security camera footage that proves it, they point to the wrong room.
- The Stat: The models' ability to find the right video segment was below 45%. They were relying on their "common sense" (what they read on the internet) rather than actually watching the video.
4. The Fix: Teaching the AI to "Look"
To help the AI learn, the researchers created a special training method (Supervised Fine-Tuning).
- The Method: They didn't just tell the AI the answer. They gave it "pseudo-labels"—fake but helpful hints showing exactly which objects (like the egg or the knife) were important and when they appeared.
- The Analogy: It's like a teacher putting a highlighter on the specific sentences in a textbook that contain the answer, rather than just giving the student the answer key.
- The Outcome: When they trained the AI with these "highlighters," the AI got much better at both answering the questions and finding the video evidence.
5. The Bonus: It Works Everywhere
The best part is that this new skill wasn't just for this specific test.
- When they took the AI trained on this "object-centric" method and tested it on other video tasks it had never seen before, it performed better than models that were just trained on general video data.
- It even helped the AI act as a "planner" for robots (embodied agents), helping them understand how to manipulate objects in the real world, not just watch videos.
Summary
The paper claims that to truly understand instructional videos, AI needs to stop just watching the "actions" and start tracking the "objects" and how they change over time. They built a new test (ProcObject-10K) to prove that current AI is "blind" to these changes, and they showed that a specific training method can fix this, making AI smarter at understanding cause-and-effect in the real world.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.