Kandinsky 5.0: A Family of Foundation Models for Image and Video Generation
This paper introduces Kandinsky 5.0, a family of open-source foundation models comprising image and video generation variants with up to 19 billion parameters, which leverage a comprehensive data curation lifecycle and advanced training techniques to achieve state-of-the-art high-resolution synthesis and 10-second video generation.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Picture: A New Family of Digital Artists
Imagine a digital art studio that has just hired a whole new family of artists. This family, called Kandinsky 5.0, is designed to create both high-definition pictures and 10-second movies just by reading your text descriptions.
The paper introduces three specific "siblings" in this family, each with a different job and skill level:
- Kandinsky 5.0 Image Lite (The 6B Artist): A specialist in creating high-resolution still images and editing existing photos.
- Kandinsky 5.0 Video Lite (The 2B Artist): A lightweight, fast worker who can make short 10-second clips. Think of them as the "speedster" of the family.
- Kandinsky 5.0 Video Pro (The 19B Artist): The heavyweight champion. This is a massive model that creates the highest quality, most detailed 10-second videos.
The main goal of this report is to show how they built these artists, how they trained them, and how they made them fast enough to be useful without needing a supercomputer in your living room.
1. The Training Camp: How They Learned to Create
You can't just give a blank canvas to a robot and expect a masterpiece. The paper details a rigorous "training camp" involving three main phases:
- Phase 1: Pre-training (The "World Tour"):
Imagine the models are students going on a massive field trip. They are shown hundreds of millions of images and videos from the internet. They learn the basics: what a cat looks like, how water moves, what a sunset feels like. This is where they learn the "grammar" of the visual world. - Phase 2: Supervised Fine-Tuning (The "Art School"):
After the field trip, they go to a strict art school. Here, human experts hand-pick the absolute best, most beautiful examples of art and video. The models practice on these perfect examples to learn how to make things look real and aesthetic, rather than just "okay." - Phase 3: The "Taste Test" (RL Post-Training):
This is a clever trick. The models generate two images. A "judge" (a reward model) looks at them and says, "I like the first one better." The model then learns to make more of what the judge likes. It's like a chef tasting their own soup and adjusting the spices until it's perfect.
2. The Secret Sauce: New Tools and Tricks
The paper explains that the old ways of building these models were too slow and clunky. They invented new tools to make things faster and sharper:
- The "Smart Attention" Mechanism (NABLA):
- The Problem: Watching a 10-second video is like trying to remember every single frame of a movie at once. For a computer, this is like trying to hold a library of books in your head while reading one page. It's too heavy.
- The Solution: They built a system called NABLA. Imagine a librarian who doesn't read every book in the library to find a quote. Instead, they only look at the specific shelves relevant to your question. This allows the model to focus only on the important parts of the video, making it 2.7 times faster without losing quality.
- The "Distillation" Process (The Speed-Up):
- The Problem: High-quality video usually takes a long time to generate (like baking a cake from scratch).
- The Solution: They created "Flash" versions of their models. Think of this as taking a master chef's recipe and creating a "microwave meal" version. It's not quite as complex as the original, but it's 90% as good and ready in a fraction of the time. They reduced the number of steps needed to make a video from 100 down to just 16.
3. The Data: Cleaning the Library
Before the models could learn, the team had to clean up the "library" of data they were learning from.
- Filtering: They threw away blurry photos, videos with watermarks, and clips that were too boring (static) or too chaotic.
- Cultural Code: They specifically gathered data related to Russian culture (architecture, history, nature) to ensure the models understand these specific visual styles, which often get missed by other global models.
- Instruction Tuning: For the image editing model, they created a massive dataset of "Before and After" pairs. If you have a photo of a dog and you want to change it to a cat, the model learned exactly how to swap the fur while keeping the background the same.
4. The Results: How Good Are They?
The team didn't just say "it's good"; they put it to the test against other famous models (like Sora, Veo, and Wan).
- The "Blind Taste Test": Humans were shown two videos side-by-side without knowing which model made them. They had to pick the winner.
- The Verdict:
- Visual Quality & Motion: Kandinsky 5.0 (especially the Pro version) often won against competitors like Veo 3 and Wan models. The videos looked more realistic, the movements were smoother, and there were fewer "glitches" (like weird melting faces).
- Following Instructions: The paper admits that while Kandinsky is great at making things look real, it sometimes struggles to follow very complex text instructions as perfectly as some competitors (like Veo 3). It's a trade-off: better visuals, slightly less precise text following.
- Speed: The "Flash" versions are incredibly fast, making them usable for real-time applications.
5. The "Open Door" Policy
Unlike some companies that keep their best models secret, the Kandinsky team is releasing everything.
- They are sharing the code, the training data recipes, and the final models for free (under an open license).
- They are not putting strict "guardrails" inside the code itself, believing that users should be responsible for how they use the tool. However, they explicitly warn that users must not use it for hate speech, deepfakes, or illegal activities.
Summary
Kandinsky 5.0 is a new, open-source family of AI artists. They use a smarter way of looking at video data to be faster, a rigorous training process to look more realistic, and a "speed-up" technique to generate clips quickly. While they are still working on perfecting their ability to follow complex text instructions, they have proven they can create stunning, high-quality images and 10-second videos that rival the best closed-source systems in the world.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.