← Latest papers
💻 computer science

Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds

The paper introduces JoyAI-Echo-1.5, a unified audio-visual generation system featuring a long-video variant with cross-shot memory for persistent character consistency and a world-model variant with geometric camera control for interactive navigation, both optimized via rollout-aware training to achieve state-of-the-art performance in long-horizon storytelling and interactive world generation.

Original authors: Nan Duan, Haoyang Huang, Weiyang Jin, Haoran Li, Yaowei Li, Yuming Li, Yijun Liu, Xin Lu, Xiaoxiao Ma, Yanwen Ma, Yaofeng Su, Yilang Sun, Haoyu Wang, Zeyue Xue, Songchun Zhang, Junhao Zhuang

Published 2026-08-25
📖 5 min read🧠 Deep dive

Original authors: Nan Duan, Haoyang Huang, Weiyang Jin, Haoran Li, Yaowei Li, Yuming Li, Yijun Liu, Xin Lu, Xiaoxiao Ma, Yanwen Ma, Yaofeng Su, Yilang Sun, Haoyu Wang, Zeyue Xue, Songchun Zhang, Junhao Zhuang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a world where a computer can not only create a single video clip but can weave an entire story, keeping characters looking and sounding the same from the first scene to the last, or even let a viewer walk through a digital landscape, turning their head and moving forward with perfect stability. This is the frontier of video generation, a field where artificial intelligence learns to synthesize moving images and sound. For years, these systems have been like skilled painters who can create a beautiful portrait but struggle to paint a whole gallery where every face remains consistent, or a storyteller who forgets the plot after a few sentences. The challenge has been to move beyond isolated moments to create long, coherent narratives and interactive worlds that do not drift apart or lose their identity as they grow.

A team of researchers at the Joy Future Academy has introduced a new system, JoyAI-Echo-1.5, designed to solve these specific problems. Rather than just making better single clips, this system focuses on two distinct but related goals: creating long-form stories where characters and voices remain consistent, and building interactive worlds that respond precisely to a user's movements. The researchers found that by giving the computer a way to "remember" past scenes and by teaching it to understand the geometry of movement, they could generate extended sequences of video that stay stable and true to the original instructions.

The first part of this system tackles the problem of long stories. In previous attempts, if a character appeared in a scene, then the video cut to a new location, the character might return looking slightly different or speaking with a changed voice. JoyAI-Echo-1.5 solves this by building a memory bank. As the system generates a story, it saves specific visual and audio details from earlier shots. When a character reappears, the system consults this memory to ensure their face, clothing, and voice match what was established before. It does this by looking at the full audio of a previous scene to capture the speaker's unique tone, rather than just a short snippet, and by saving images of the character from different angles. This allows the computer to generate a sequence of shots that feel like a continuous movie, where a hero can travel through different environments without losing their identity.

The second part of the system is designed for interactive worlds, where a user might want to control a camera to explore a digital environment. In the past, giving a computer a command like "move forward" or "turn left" often resulted in jerky, unnatural movement or a scene that slowly warped and changed shape as the video played on. The new system translates these commands into a precise, mathematical description of how the camera moves through space. It treats the world as a stable, three-dimensional space that exists independently of the camera. When a user moves, the system calculates the exact path and adjusts the video accordingly, ensuring that the walls, objects, and lighting remain consistent no matter how long the exploration continues. This allows for a smooth, continuous experience where the digital world feels solid and real, rather than a shifting illusion.

To make these systems fast enough to be useful, the researchers also developed a new way to train them. Usually, creating high-quality video requires many steps of calculation, which takes a long time. The team taught the system to learn from its own mistakes by having it generate video, then immediately trying to improve that video based on what it just created. This process, which involves a technique called self-gradient forcing, allows the system to become stable even when it is generating long sequences on its own. They also used a method to compress the generation process, allowing the system to produce high-quality video in just a few steps instead of dozens, without losing the fine details or the synchronization between sound and image.

The results of this work were tested against other leading systems. In tests involving long stories, the new system was better at keeping characters consistent and matching the spoken words to the lip movements than any previous model. In tests of interactive worlds, it scored higher than its competitors in maintaining the shape of the environment and following the user's movement commands accurately. One specific test showed that the system could generate a sixty-second continuous video where the camera moved through a complex scene, and the system maintained the correct perspective and visual quality throughout, a feat that other models struggled to achieve without the scene drifting or distorting.

This research suggests that the key to creating persistent stories and stable interactive worlds lies in how the computer remembers the past and understands the geometry of the present. By combining a robust memory for visual and audio details with a precise understanding of movement, the system can generate content that feels coherent over time. While challenges remain, particularly in keeping the camera perfectly steady during very long and complex movements, the work demonstrates a significant step forward. It shows that artificial intelligence can move beyond creating isolated clips to sustaining entire narratives and evolving worlds, opening the door for more immersive and reliable digital experiences.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →