← Latest papers
💻 computer science

Perception Test 2025: Challenge Summary and a Unified VQA Extension

This paper summarizes the results of the Perception Test 2025 challenge, which evaluated state-of-the-art multimodal video models on a unified benchmark featuring consolidated tracks like unified video QA and tracking to highlight the significant difficulties current models face when addressing diverse perception tasks through a single interface.

Original authors: Joseph Heyward, Nikhil Parthasarathy, Tyler Zhu, Aravindh Mahendran, João Carreira, Dima Damen, Andrew Zisserman, Viorica Pătrăucean

Published 2026-04-30
📖 4 min read☕ Coffee break read

Original authors: Joseph Heyward, Nikhil Parthasarathy, Tyler Zhu, Aravindh Mahendran, João Carreira, Dima Damen, Andrew Zisserman, Viorica Pătrăucean

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a giant, high-stakes talent show for computer vision, but instead of singing or dancing, the contestants are artificial intelligence models trying to "see" and understand the world. This paper is the report card from the 2025 edition of this show, called the Perception Test, held alongside a major computer science conference.

Here is the breakdown of what happened, using simple analogies.

The Big Idea: From Specialized Tools to a Swiss Army Knife

In the past, AI models were like specialized tools: one model was great at finding a specific person in a crowd (tracking), another was great at spotting when a car crash happened (action detection), and a third was good at answering questions about a video.

For this year's challenge, the organizers decided to stop testing these "specialized tools." Instead, they demanded a Swiss Army Knife. They wanted to see if a single AI model could do everything at once without switching hats. They asked: "Can one brain handle tracking a ball, hearing a sound, and answering a question about the scene all at the same time?"

The Five Main Events (Tracks)

The competition had five main events, each testing a different "muscle" of the AI:

  1. The Unified Video Quiz (The "Magic Eye" Test):

    • The Old Way: To test if an AI could track a moving dot, you had to build a specific tracker.
    • The New Way: The organizers turned these hard tasks into multiple-choice questions. Imagine a video where a specific spot on a table flashes green. The AI is asked, "Which of these five dots was the one that flashed?"
    • The Twist: They added tricky questions like "What order did these actions happen?" or "Which object was pretending to do something?" This forced the AI to use language to solve visual puzzles.
  2. The Double-Tracking Challenge:

    • The AI had to follow two things at once: a whole object (like a bouncing ball) and a specific point on that object (like a red sticker on the ball), even if the ball went behind a wall (occlusion).
  3. The Sound-and-Action Detective:

    • The AI had to watch a video and say, "At exactly 10 seconds, someone kicked a ball, and you heard a thud." It had to find the when and what for both the visual action and the sound simultaneously.
  4. The "Point-and-Click" Detective:

    • The AI was asked a question like, "Which dog is chasing the cat?" and had to not just say "The brown one," but actually draw a box around that specific dog and follow it through the whole video.
  5. The Marathon Runner (Hour-Long Videos):

    • Most AI models get tired after watching a 30-second clip. This event threw 1-hour videos at them. The AI had to remember details from the beginning to answer a question at the end.

The Results: A Tale of Two Speeds

The paper reports a fascinating split in how well the AI did:

  • The Language Winners: When the AI could use language (answering questions, describing scenes), it got much better than last year. In fact, for the "Point-and-Click" and "Hour-Long" tests, the scores nearly doubled. It's as if the AI suddenly learned to read a manual and apply it to the visual world.
  • The "Silent" Struggle: When the AI had to do things without language (like just tracking a dot or finding a sound), it didn't improve much. The best "Swiss Army Knife" models this year were actually slower than the specialized "single-task" models from last year.

The Takeaway: The paper concludes that while AI is getting amazing at talking about what it sees, it is still struggling to be a single, unified brain that can do everything perfectly at once. The "general perception" model is on the horizon, but it's not quite there yet.

The Winners

  • Total Submissions: Over 100 teams sent in 450 attempts.
  • The Prize: The winners took home a total of $47,000.
  • Top Performers: The team "NJUST_KMG" was a frequent winner, taking top spots in the tracking, sound, and hour-long video categories. Another team, "SGVR@KAIST," won the "Point-and-Click" category.

In a Nutshell

The 2025 Perception Test showed us that AI is rapidly learning to talk about what it sees, but it is still learning how to do everything at once without getting confused. The gap between "specialized robots" and "general human-like perception" is shrinking, but the finish line is still a bit away.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →