← Latest papers
💻 computer science

KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation

This paper introduces KeyFrame-Compass, the first comprehensive benchmark and automated evaluation framework for keyframe-conditioned video generation, which reveals that current models struggle to balance faithful keyframe execution with natural video synthesis, particularly under dense constraints or storyboard-grid inputs.

Original authors: Yuqi Tang, Tengfei Liu, Yizheng Lai, Yuran Wang, Yang Shi, Wanshun Su, Zhuoran Zhang, Qixun Wang, Xiaohan Zhang, Xinlei Yu, Xuehai Bai, Xuanyu Zhu, Bohan Zeng, Bozhou Li, Shujie Li, Yifan Dai, Yujie W
Published 2026-07-17
📖 4 min read☕ Coffee break read

Original authors: Yuqi Tang, Tengfei Liu, Yizheng Lai, Yuran Wang, Yang Shi, Wanshun Su, Zhuoran Zhang, Qixun Wang, Xiaohan Zhang, Xinlei Yu, Xuehai Bai, Xuanyu Zhu, Bohan Zeng, Bozhou Li, Shujie Li, Yifan Dai, Yujie Wei, Shixuan Liu, Haotian Wang, Jialu Chen, Yuanxing Zhang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are the director of a movie, but instead of hiring actors and building sets, you are talking to a magical robot that can conjure moving pictures out of thin air. For a long time, these robots were great at making short clips based on a single picture or a simple sentence like "a cat running." But now, creators want to tell full stories. They want to give the robot a sequence of specific drawings—like a comic book or a storyboard—and say, "Make this happen, in this order, with these exact characters." This is the world of keyframe-conditioned video generation. Think of "keyframes" as the anchor points of a story: the starting pose, the middle action, and the final landing. The challenge is getting the robot to not just copy those pictures, but to fill in the smooth, natural movement between them without messing up the characters or the timeline. If the robot gets it wrong, you might end up with a video where the hero suddenly turns into a villain, or the scenes jump around like a broken DVD.

Enter KeyFrame-Compass, a new "report card" designed to grade how well these video-making robots are actually doing their job. Created by a team of researchers, this isn't just another test that asks, "Does this video look pretty?" Instead, it acts like a strict film critic who checks the script line-by-line. They built a library of 386 specific test cases, ranging from everyday moments like "a dog chasing a ball" to complex movie scenes and product ads. They tested nine different video-generation systems, including both famous commercial ones and open-source projects, to see if they could faithfully reproduce a sequence of images in the right order and at the right time.

The results? It's a bit of a mixed bag, revealing a tricky balancing act. The researchers found that most models struggle to be both a "loyal copycat" and a "creative storyteller" at the same time. Some models, like LTX-2.3, are incredibly good at sticking to the pictures you give them—they hit the right frames with high accuracy. However, when they try to connect the dots between those frames, the movement often feels jerky, like a slideshow where the pictures just snap together instead of flowing smoothly. On the other hand, models like Gemini-Omni-Flash are great at making smooth, high-quality movies that look natural, but they often ignore the specific pictures you gave them, inventing new characters or scenes instead of following your storyboard.

The study suggests that as you give the robots more pictures to follow (increasing the "density" of constraints), they get worse at following instructions. It's like trying to follow a recipe with more and more steps; the more specific you get, the more likely the chef is to drop a spoon or forget an ingredient. Furthermore, the paper highlights a major gap between "closed" (proprietary) models and "open" (public) ones. While the top commercial models can understand a grid of storyboard images and turn them into a movie, many open-source models fail completely. When shown a grid of pictures, they often just copy the grid itself as a static image, rather than understanding that each square is a different moment in time.

Ultimately, KeyFrame-Compass suggests that while we are getting closer to robots that can direct movies, they still have a long way to go. They haven't quite mastered the art of connecting specific visual anchors with natural, physics-defying motion, and they often struggle to understand complex instructions when the task gets too crowded. The paper doesn't claim to have solved the problem, but it provides the first clear map of exactly where the robots are stumbling, helping future developers know whether to focus on making the robot stick to the script or on making the movie look smoother.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →