Results of the 1st Asynchronous CASTLE Challenge at the Joint Egocentric Vision Workshop in Conjunction with CVPR 2026
This paper summarizes the contributions and results of the 1st Asynchronous CASTLE Challenge, which was held at the Joint Egocentric Vision Workshop in conjunction with CVPR 2026.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine trying to remember everything that happened during a four-day vacation, not just the highlights, but every conversation, every object moved, and every small change in the room, all while watching from the eyes of ten different people simultaneously. This is the kind of memory test that computers struggle with, even when they are fed thousands of hours of video. In the field of artificial intelligence, researchers are trying to build systems that can watch long, unscripted recordings of daily life and answer specific questions about them, much like a human would after reading a diary or watching a movie. The challenge lies in the sheer volume of information and the need to connect details that might be separated by hours or days. If a machine can learn to do this, it could eventually help us search through years of personal footage to find a specific moment or understand complex patterns in how people live.
To test how close we are to solving this, a group of researchers organized a competition called the CASTLE Challenge. They provided teams with a massive dataset containing over 600 hours of ultra-high-definition video, recorded at a high speed to capture every detail. The footage came from cameras worn by ten participants as they went about their normal lives, along with five stationary cameras watching from the room. The data was rich, including audio, heart rate monitors, and even thermal images, creating a complete digital record of four days of activity. The task for the competing teams was simple in theory but incredibly difficult in practice: they had to answer 185 multiple-choice questions about what happened in the videos. These questions required the computer to track people, count objects, remember where things were hidden, and understand the timing of events, such as figuring out who ate a slice of cake first or how fast someone was charging their car on the final day.
Four teams stepped up to the challenge, each bringing a different strategy to the problem. The winning team, known as WDL, treated the task like a detective gathering clues. Instead of trying to watch all 600 hours of video at once, their system first looked at the question to find keywords, such as a person's name or a specific day. It then used these clues to pull out only the most relevant parts of the video and the spoken transcripts, ignoring the rest. They trained their computer model to focus strictly on the evidence it found, running strict, balanced, and aggressive prompt variants for each question to ensure consistency through multi-sample self-consistency. This approach allowed them to reach an accuracy of nearly 58 percent, meaning they got about 108 of the 185 questions right.
Another team, MARS, took a more interactive approach. They built a system that acted like a curious agent, deciding step-by-step which type of information to look at next. If a question was about physical activity, the system might choose to check the heart rate data specifically for those queries. If it was about where someone was looking, it would pull up the eye-tracking data. By dynamically choosing which piece of evidence to examine rather than trying to process everything at once, they improved their score to 57 percent. A third team, TAHAKOM, organized the information into a structured map of relationships, connecting people, places, and objects with time stamps. This allowed them to ask the system specific questions about the connections between events, achieving an accuracy of 54.6 percent. The final team, CuriosAI, focused on preventing the computer from making things up. They built a strict process where the system had to search for evidence, verify it against the raw video, and then answer, ensuring that every claim was grounded in what was actually seen or heard. Their best method reached 49.7 percent accuracy.
The results of the competition show that while computers are getting better at understanding long videos, they still have a long way to go. The top teams answered more than half of the questions correctly, which is a significant improvement over random guessing, but they still missed nearly half of the answers. What is perhaps most revealing is that the teams did not all get the same questions right. When you combine the correct answers from all four teams, they managed to solve 176 of the 185 questions, leaving only nine that no team could answer. This suggests that no single method is perfect yet, but different approaches capture different pieces of the puzzle. The researchers found that the biggest hurdle remains the ability to cover enough evidence without getting lost in the sheer volume of data. While the systems can navigate hundreds of hours of footage, the task of reasoning through unscripted, real-world events over long periods remains a difficult problem. The competition did not solve the issue, but it clearly showed that combining these different strategies could be the key to building machines that truly understand our daily lives.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.