← Latest papers
💻 computer science

MUSE: A Multi-agent Framework for Unconstrained Story Envisioning via Closed-Loop Cognitive Orchestration

This paper introduces MUSE, a multi-agent framework that addresses the challenges of generating long-form audio-visual stories by employing a closed-loop cognitive orchestration process to iteratively plan, execute, verify, and revise content, thereby ensuring narrative coherence and identity consistency across extended sequences.

Original authors: Wenzhang Sun, Zhenyu Wang, Zhangchi Hu, Chunfeng Wang, Hao Li, Wei Chen

Published 2026-02-04
📖 5 min read🧠 Deep dive

Original authors: Wenzhang Sun, Zhenyu Wang, Zhangchi Hu, Chunfeng Wang, Hao Li, Wei Chen

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you want to tell a long, complex story to a friend, but you only have a single sentence to start with: "A boy goes on an adventure in a forest."

If you ask a standard AI to turn that sentence into a movie, it might make a beautiful first scene. But by the time it gets to the fifth or tenth scene, the boy's face might look like a different person, the forest might suddenly turn into a desert, or the story might forget what happened in the first scene. The AI gets "drunk" on the immediate task of making the next picture, losing the big picture of the whole story.

MUSE is a new system designed to fix this. Think of MUSE not as a single artist, but as a highly organized film production crew working together to ensure the story stays consistent from the first frame to the last.

Here is how MUSE works, broken down into simple parts:

1. The Problem: The "Drunk Artist"

Current AI video generators are like a talented but short-sighted artist. They are great at painting one beautiful picture based on a prompt. But if you ask them to paint a whole movie, they forget the rules. The main character might change clothes, the lighting might flip from day to night randomly, or the character's voice might sound like a different person in the next scene. This is called "semantic drift."

2. The Solution: A "Closed-Loop" Film Crew

MUSE solves this by treating storytelling like a closed-loop factory rather than a one-way street. It uses a team of specialized AI agents (digital workers) that constantly check and correct each other.

The process happens in three main stages, like a real movie production:

  • Pre-Production (The Casting Director & Scriptwriter):
    Before any video is made, MUSE creates a "Character Bible." It locks in exactly what the main character looks like (age, clothes, style) and how they sound (voice pitch, tone). It creates a digital "ID card" for the character. This ID card is the rulebook that must be followed for the entire movie.

    • Analogy: Imagine a casting director who takes a perfect photo and a voice recording of an actor and says, "This is who we are using. Do not change a thing."
  • Production (The Director & Set Designer):
    Now, the system starts generating the scenes. But it doesn't just guess. It uses the "ID card" to ensure the character looks the same. If the script says the character is holding a sword, MUSE checks the layout to make sure the sword is actually there and not floating in mid-air.

    • Analogy: The director is constantly shouting, "Wait! That actor looks different! Fix it!" before the camera even rolls.
  • Post-Production (The Editor & Continuity Supervisor):
    Once the scenes are made, MUSE checks the "seams" between them. Did the character's hand move smoothly from the last shot to this one? Did the story flow logically? If the AI accidentally made the character walk through a wall or if the voice sounds too happy for a sad scene, MUSE catches it.

    • Analogy: An editor who watches the movie and says, "This scene doesn't match the last one. Let's fix the transition before we show it to the audience."

3. The "Plan-Execute-Verify-Revise" Loop

The secret sauce of MUSE is that it doesn't just generate and hope for the best. It follows a strict cycle:

  1. Plan: Decide what needs to happen next.
  2. Execute: Generate the video/audio.
  3. Verify: A "critic" agent checks if the result matches the plan and the character ID.
  4. Revise: If something is wrong (e.g., the character's hat is missing), MUSE fixes only that specific part, rather than throwing away the whole scene and starting over.

4. The New Test: MUSEBench

How do you know if a story is good if there is no "correct" answer? The authors created a new testing ground called MUSEBench.
Instead of comparing the AI's output to a perfect reference video (which doesn't exist for creative stories), MUSEBench uses a smart AI judge to grade the story on things like:

  • Consistency: Did the character stay the same?
  • Story Flow: Did the plot make sense?
  • Audio-Visual Harmony: Did the voice match the mood of the scene?

The Result

The paper shows that MUSE is much better at keeping long stories coherent than previous methods. It keeps characters looking and sounding the same, maintains the right mood, and ensures the story doesn't fall apart as it gets longer.

In short: MUSE is like a super-organized film crew that refuses to let the story get messy. It plans ahead, checks its work constantly, and fixes mistakes immediately, ensuring that a simple prompt can turn into a long, immersive, and consistent movie.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →