← Latest papers
🤖 AI

S-EMBER: A Large-Scale Benchmark for Streaming Egocentric Memory Retrieval

The paper introduces S-EMBER, a large-scale benchmark of 388 hours of smart-glasses footage designed to evaluate AI agents' ability to perform causal, streaming episodic memory retrieval, revealing that while semantic reasoning scales with model size, precise temporal grounding remains a persistent architectural bottleneck.

Original authors: Xiaodong Wang, Xuanyi Zhao, Pedro Rodriguez, Devendra Singh Sachan, Barlas Oguz, Seungwhan Moon, Shang-Wen Li, Gargi Ghosh, Xin Dong, Wen-Tau Yih

Published 2026-07-07
📖 5 min read🧠 Deep dive

Original authors: Xiaodong Wang, Xuanyi Zhao, Pedro Rodriguez, Devendra Singh Sachan, Barlas Oguz, Seungwhan Moon, Shang-Wen Li, Gargi Ghosh, Xin Dong, Wen-Tau Yih

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a smart assistant that lives in your glasses. It's been recording your entire day, from the moment you wake up to the moment you go to bed. Now, imagine you ask it, "What ingredients did I buy for those cookies I baked last Tuesday, and where did I get them?"

To answer this, the assistant can't just guess or make things up. It has to dig through hours of video footage, find the exact moment you were at the grocery store, and tell you the truth. This is the challenge the paper S-EMBER tackles.

Here is a breakdown of the paper using simple analogies:

1. The Problem: The "Library" vs. The "Live Stream"

Most AI tests today are like giving a student a whole book and asking them a question about page 50. The student can flip back and forth, read the whole thing, and find the answer. This is called "offline" testing.

But real life isn't a book you can flip through. It's a live stream. Your AI assistant only sees what is happening right now and what happened just before. It can't peek into the future or rewind the tape to see the whole movie at once. The paper argues that current AI models are great at reading the whole book but terrible at navigating a live stream.

2. The Solution: S-EMBER (The "Memory Gym")

The authors created a massive new test called S-EMBER. Think of it as a gym for AI memory.

  • The Workout: They recorded 388 hours of real life using Ray-Ban Meta smart glasses worn by over 600 different people.
  • The Exercises: They created nearly 10,000 questions based on these videos. These aren't simple questions like "What color is the car?" They are "memory" questions like, "How long did I spend chopping onions?" or "Where did I put my keys before I left the house?"
  • The Twist: Every question comes with a "proof" requirement. The AI doesn't just have to say the answer; it has to point to the exact time interval in the video where the answer is hiding. It's like asking a detective not just to solve the crime, but to show the exact minute on the security camera where the suspect was seen.

3. The Big Discovery: The "Localization Paradox"

This is the most surprising part of the paper. The researchers tested the world's smartest AI models on this gym. They found a strange contradiction, which they call a paradox:

  • The Brain is Getting Bigger: As AI models get bigger and smarter (more "parameters"), they get much better at understanding what is happening. If you ask, "What is this person doing?", a bigger model answers correctly more often.
  • The Clock is Broken: However, when it comes to when it happened, the models hit a wall. No matter how big the model gets, or how many video frames they feed it, they are terrible at pinpointing the exact time.

The Analogy: Imagine a detective who is a genius at describing a suspect's face and clothes (Semantic Reasoning) but is completely blind to the clock on the wall (Temporal Grounding). Even if you give the detective a super-powerful brain or a high-definition camera, they still can't tell you what time the crime happened. They just guess.

4. The "Needle in a Haystack" Problem

The videos are full of boring, everyday stuff (making coffee, walking the dog). The answer to a question might be hidden for just a few seconds in the middle of a 20-minute video.

  • The Result: The AI models are like people searching for a needle in a haystack. If you give them more hay (more video frames), they get slightly better at finding the needle. But they still struggle to say exactly where in the haystack the needle is buried.
  • The "Recency" Effect: The study also found that AI memory fades fast. If a question is about something that happened 10 minutes ago, the AI's accuracy drops significantly. It's like a human who can remember what they had for breakfast but forgets what they did an hour ago.

5. Why This Matters

The paper concludes that we can't just make AI models bigger to fix this. We need to change how they are built. Currently, they are "brute-forcing" the problem by looking at more data, but they lack a specific architectural tool to track time accurately.

In short: We have built a massive test to see if AI can remember our daily lives like a human does. The test shows that while AI is getting smarter at understanding what we do, it is still failing at remembering when we did it. Until we fix this "time blindness," our AI assistants won't be reliable enough to help us navigate our real-world memories.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →