← Latest papers
🔬 optics

Engram-E2VID: Reference-Based Event-to-Video Reconstruction via Generative Activation of Appearance Engrams

Engram-E2VID is a reference-based event-to-video reconstruction framework that leverages a one-step diffusion backbone to structurally guide the generative activation of appearance engrams encoded from a reference frame, enabling high-fidelity recovery of target RGB frames from sparse event streams even under complex motion and long temporal intervals.

Original authors: Feiyu Ji, Xiang Li, Hao Ma, Tianxiang Huang, Qingxin Lu, Mengqi Ji, Lei Han, Xiaokang Yang, Xiaoyun Yuan

Published 2026-08-07
📖 4 min read☕ Coffee break read

Original authors: Feiyu Ji, Xiang Li, Hao Ma, Tianxiang Huang, Qingxin Lu, Mengqi Ji, Lei Han, Xiaokang Yang, Xiaoyun Yuan

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to recreate a movie scene, but you only have two very strange tools. The first is a single, perfect photograph of the scene before the action starts. The second is a chaotic stream of tiny, invisible sparks that fly around whenever something moves or changes brightness. These sparks, called "events," are incredibly fast and detailed about movement, but they are completely blind to color, texture, or what things actually look like. They are like a ghostly whisper of motion without a body.

Scientists have long tried to use these sparks to rebuild the missing parts of a movie, but it's like trying to paint a portrait using only a map of where the wind blew. The old methods often get lost when things move too far or too fast, resulting in blurry, ghostly, or completely wrong pictures. This is the puzzle that a team of researchers at Shanghai Jiao Tong University and other institutions decided to solve. They asked: "How can we take that chaotic stream of motion sparks and combine it with our starting photo to rebuild a crystal-clear, new frame of the video, even if a lot of time has passed?"

The answer they found is a new system called Engram-E2VID. Think of it as a magical construction crew that doesn't just copy-paste the old photo. Instead, it builds a "scaffold" or a skeleton of the new scene using the motion sparks. This scaffold tells the system where things are moving and how the structure is changing. Then, the system reaches into its memory bank—the original photo—and pulls out the specific "appearance engrams" (which are like detailed memory fragments of how the objects look).

Here is the clever part: instead of trying to stretch the old photo to fit the new position (which usually makes it look warped and blurry), the system uses a smart "activation" process. It asks the scaffold: "Hey, I need the texture of a red ball here, and the fur of a cat there." The scaffold points to the right spots, and the system grabs the perfect memory fragments to fill them in. If there are parts of the scene that were completely hidden or new, a powerful AI "imagination" (called a diffusion model) steps in to fill in the blanks realistically.

The results are impressive. When tested on three different real-world datasets, this new method didn't just work; it worked much better than the previous best attempts. In simple terms, it made the reconstructed images significantly sharper and more accurate. Specifically, it improved the picture quality score (PSNR) by up to 3.29 dB and reduced visual fuzziness (LPIPS) by up to 0.08 compared to the strongest existing method that used the same inputs.

Perhaps the most exciting finding is how well it handles time. Usually, the longer you wait between the starting photo and the new frame you want to create, the worse the picture gets. But Engram-E2VID degrades much more slowly. Even when the time gap was large, it kept the details sharp and the structures correct, whereas other methods started to blur and hallucinate weird shapes. The researchers suggest that by separating the "structure" (the scaffold) from the "appearance" (the engrams) and letting them talk to each other in a smart way, they can reconstruct scenes that are complex, moving fast, or far away from the original photo, all without needing a second photo to help them out. It's a bit like being able to remember exactly what a friend's face looks like, even if they have moved to a completely different room and are running around, just by knowing the layout of the room and the sound of their footsteps.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →