MEval: Multi-Modal Memory Evaluation through Cognitively-Grounded Video Tasks
This paper introduces MEval, the first cognitively-grounded benchmark designed to systematically evaluate memory capabilities in multi-modal models through specialized video tasks, revealing critical weaknesses such as poor disentanglement in parallel streams and limited symbolic memory that highlight the need for improved memory mechanisms.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a new student who is incredibly smart at looking at pictures and reading text. They can describe a single movie scene perfectly. But now, you want to see if they can actually remember what happened in a whole movie, especially when things get complicated, like when two movies are playing at once or when similar scenes are mixed together.
The paper "M 3Eval" is like a new, very specific report card designed just for this. The researchers realized that while we have tests for how well AI "sees" and "thinks," we don't have good tests for how well it "remembers." So, they built a gym for AI memory, based on how human psychologists test our own brains.
Here is how they tested the AI, using four simple games:
1. The "Split-Screen" Test (Divided Attention)
The Setup: Imagine watching two different cooking shows at the same time, side-by-side on your TV. One is on the left, one is on the right. Every few seconds, the TV swaps them so the left one moves to the right and vice versa.
The Question: "In the video that was originally on the left, did the chef add onions or tomatoes?"
The Result: The AI got very confused. It couldn't keep the two stories separate. It was like a person trying to listen to two different conversations at a loud party; the AI started mixing up who said what and where. When the videos swapped places, the AI lost track of which story belonged to which screen.
2. The "Interference" Test (Memory Interference)
The Setup: Imagine you watch a video of a man baking a cake (Video A). Immediately after, you watch a very similar video of a man baking a pie (Video B).
The Question: "What did the man in the first video (the cake) put on top?"
The Result: The AI struggled to keep the two memories apart. It often answered with details from the pie video instead of the cake video.
The Twist: The researchers found something surprising. If they showed the cake video twice before showing the pie, the AI remembered the cake much better. It's like if you hear a song twice, you're less likely to forget it when a similar song plays right after.
3. The "Scrambled Story" Test (Interleaved Events)
The Setup: Imagine taking two different movies, cutting them into 10 tiny pieces each, and then shuffling them together like a deck of cards (Piece 1 of Movie A, Piece 1 of Movie B, Piece 2 of Movie A, Piece 2 of Movie B...).
The Question: "Can you tell me the correct order of events for just Movie A?"
The Result: The AI was terrible at this. It couldn't untangle the two stories. It was like trying to reassemble two different jigsaw puzzles that have been mixed into one big pile. The AI also had a hard time knowing when something happened (temporal order) compared to where it happened (spatial location). It was better at remembering "The pot was on the left" than "The pot was added before the onions."
4. The "N-Back" Test (Symbolic Memory)
The Setup: This is a classic brain game. You watch a long stream of short video clips. The rule is: "If the current clip looks like the one you saw 3 clips ago, say 'Yes'."
The Result: Humans get worse at this as the gap gets bigger (it's hard to remember 10 steps back). But the AI was weird: it didn't get worse as the gap got bigger, but it got terrible as the total number of clips got longer.
The Metaphor: Imagine a human has a small notepad and can only write down the last few things they saw. If you show them too many things, they forget the old ones. The AI, however, has a giant library where it keeps everything, but it doesn't know how to ignore the junk. It gets overwhelmed by the sheer volume of information, not by the distance in time.
The Big Takeaway
The paper concludes that while AI is getting better at understanding video, its "memory" is still very different from a human's.
- Humans are good at filtering out noise and remembering the order of events, even if we forget details over time.
- AI struggles to keep parallel stories separate, gets confused by similar-looking videos, and has a hard time organizing a long list of events in the right order.
The researchers hope that by using these specific "games," other scientists can build better AI systems that don't just "see" the world, but actually remember it the way we do. They have made their test questions and videos available for everyone to use.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.