SF20K Competition 2025: Summary and findings
The first SF20K Competition at ICCV 2025 evaluated 22 teams on open-ended short-film question answering, revealing that narrative-aware processing and effective information selection are more critical than raw model capacity, though a significant performance gap remains between current AI systems and human comprehension.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you're hosting a massive movie night, but instead of watching famous blockbusters like The Avengers or Star Wars, you're watching 20,000 homemade short films made by regular people. These films are about 12 minutes long and tell unique, complex stories.
Now, imagine you want to test how smart a computer is at understanding these stories. You can't just ask, "Who is the hero?" because the computer might just guess based on what it knows about famous movies. Instead, you ask specific questions about the plot, like, "Why did the character in the blue shirt leave the room at the 8-minute mark?"
This is exactly what the SF20K Competition was all about. It was a contest held in 2025 where 22 teams of AI researchers tried to build computers that could watch these amateur films and answer questions about the story.
Here is a breakdown of what happened, using some simple analogies:
1. The Challenge: Reading a Book vs. Skimming a Menu
Most AI today is really good at looking at short clips (like a 5-second TikTok) and saying, "That's a dog running." But this competition asked the AI to do something much harder: read the whole book.
The AI had to remember characters, track cause-and-effect (e.g., "He dropped the cup, so it broke"), and understand the flow of a story over 12 minutes. The goal was to see if the AI could actually watch and understand, or if it was just memorizing answers from its training data.
2. The Rules of the Game
The competition had two main categories, like a "Pro Division" and a "Junior Division":
- Main Track: Teams could use any size computer brain they wanted (even the biggest, most expensive ones).
- Special Track: Teams had to use "small" brains (under 8 billion parameters). This was to see if clever tricks could beat raw power.
The test was strict. The AI couldn't just guess. If it didn't watch the movie, it would fail miserably (scoring only 14.7% on a "blind guess" test).
3. The Winners: It's Not About Muscle, It's About Strategy
The big surprise wasn't who had the biggest computer. It was how they used it.
- The "Brute Force" Approach: Some teams tried to feed the AI every single frame of the movie (like trying to read a book by staring at every single letter at once). This didn't work well.
- The "Smart Editor" Approach: The winning team (WXYZ) treated the movie like a storybook with chapters. They didn't look at every frame. Instead, they:
- Cut the movie into "shots" (like scenes in a play).
- Listened to the audio to find the rhythm of the story (when things got exciting or quiet).
- Summarized each scene before answering the question.
The Analogy: Imagine you have to write a book report.
- Team A tries to read every single word of the book 10 times. They get tired and miss the point.
- Team B reads the table of contents, summarizes each chapter, and then writes the report.
- Result: Team B (the smaller, smarter team) got a better grade than Team A (the giant, brute-force team).
In fact, the winning team used a relatively small AI model, but because they organized the information so well, they beat a model that was 30 times larger but used a simpler method.
4. The Secret Ingredient: The Subtitles
The researchers found that what the characters said was often more important than what they looked like.
- The AI struggled if it only watched the video.
- The AI did much better when it had a transcript (subtitles) of what was being said.
- Teams that spent extra time making better subtitles (fixing errors or using better speech-to-text tools) got significantly higher scores. It's like trying to understand a foreign movie: if the subtitles are bad, you miss the plot; if they are perfect, you get it.
5. The Final Score: Humans Still Win
Even though the AI teams did a great job, there is still a huge gap between them and humans.
- Human Score: 91.7% (Humans are great at understanding stories).
- Best AI Score: 65.7% (The best computer got about two-thirds right).
- Small AI Score: 48.7% (The "Junior Division" winner).
The Big Takeaway
The main lesson from this competition is that being smart isn't just about having a bigger brain.
For computers to understand long stories, we don't just need to make them bigger. We need to teach them how to organize information, how to pick out the important parts (like the key scenes), and how to use the dialogue to make sense of the visuals. The "bottleneck" isn't the size of the computer; it's the strategy we use to feed it the story.
Right now, AI is like a very fast reader who can memorize facts but still struggles to understand the feeling and flow of a complex story. We are getting there, but we still have a long way to go to reach human-level understanding.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.