← Latest papers
🤖 AI

ViSAGE: Constructing Self-Correcting Memories for Long-Form Video Understanding

ViSAGE is a multimodal agentic memory framework that enhances long-form video understanding by constructing self-correcting, entity-centric memories through cross-modal binding, bidirectional refinement, and multi-agent cross-verification to prevent identity confusion and hallucinations, achieving a 5.9% accuracy improvement over state-of-the-art baselines.

Original authors: Xinkui Zhao, Enbo Chen, Yifan Zhang, Chang Liu, Guanjie Cheng, Naibo Wang, Yueshen Xu

Published 2026-08-03
📖 4 min read☕ Coffee break read

Original authors: Xinkui Zhao, Enbo Chen, Yifan Zhang, Chang Liu, Guanjie Cheng, Naibo Wang, Yueshen Xu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to solve a mystery that spans an entire movie, but you are only allowed to look at one frozen frame at a time, and every few minutes, your memory gets wiped clean. This is the daily struggle of modern artificial intelligence when it tries to understand long videos. Today's smartest AI models are like brilliant detectives who can read a single page of a book perfectly, but if you hand them a 40-minute film, they get overwhelmed. They try to squish the whole story into a tiny mental box, or they chop the movie into tiny, isolated pieces. The problem? When you chop a story, you lose the connections. You forget who the characters are, you mix up their names, and you start guessing wildly about what happened in the first scene because you can't remember the clues from the last scene.

This is where the field of "long-form video understanding" comes in. Scientists are trying to teach computers to watch hours of video, remember the details, and answer tricky questions about what happened, who did it, and why. The goal is to build an AI that doesn't just see pixels but understands a story, keeping track of people and objects from the opening scene to the credits. But current methods often fail because they are too aggressive in summarizing the video, throwing away the tiny details needed to tell one person from another, or they get stuck in a loop where they can't fix a mistake made five minutes ago because they are only looking at the "now."

Enter ViSAGE, a new system designed by researchers at Zhejiang University and Xidian University that acts like a super-powered, self-correcting memory for AI agents. Think of standard AI memory as a messy notebook where you scribble notes as you watch a movie, but once you turn the page, you can't go back to fix a wrong guess. If you think a character is "Mario" in the first scene, but later find out it's actually "Emma," a normal system keeps calling them Mario forever, leading to confused and wrong answers. ViSAGE, however, is like a detective who keeps a master list of suspects and a timeline of events. When a new clue arrives—like hearing a name mentioned in a conversation at the 30-minute mark—it doesn't just file it away; it goes back and rewrites the notes from the beginning to make sure the whole story makes sense.

The researchers found that by building this "self-correcting" system, the AI could finally handle the confusion of long videos. They created a framework that separates the "what happened" (the events) from the "who did it" (the characters). When the AI sees a person but doesn't know their name yet, it gives them a temporary label. Later, when the video reveals their name, ViSAGE uses a "backward update" to instantly fix all the previous notes, ensuring that every action is correctly linked to the right person. It also uses a team of "agents" to double-check its work. If the AI isn't sure about an answer, instead of making up a story (which is called hallucinating), it is programmed to say, "I don't know," which is much safer and more reliable.

In their tests, this new approach worked significantly better than the previous best methods. On a set of challenging long-video tests, ViSAGE improved accuracy by 5.9% compared to the strongest existing system. Specifically, it scored 45.5% on robot-related video tasks, 58.4% on web-based video tasks, and a massive 79.1% on the Video-MME-long benchmark. The researchers showed that by fixing the memory errors, the AI didn't just get more questions right; it also stopped making dangerous mistakes, like mixing up who was doing what, and became much better at admitting when it lacked information. This suggests that for AI to truly understand long stories, it needs a memory that can grow, change, and correct itself, rather than just a static list of facts.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →