← Latest papers
💬 NLP

From Sports to Safety: Benchmarking Proactive Risk Inference in MLLMs

This paper introduces SPRINT, a comprehensive benchmark of real-world sports videos designed to evaluate the proactive risk inference capabilities of Multimodal Large Language Models (MLLMs), revealing that while current models excel at detecting hazards, they struggle to identify their underlying causes and frequently generate false alarms, highlighting a critical gap in reliable proactive safety for dynamic physical environments.

Original authors: Jiawei Qiu, Yichen Xu, Jianzhe Ma, Mingyang Yu, Wenbin Zhu, Yang Han, Pinzheng Lv, Wenxuan Wang

Published 2026-08-07
📖 4 min read☕ Coffee break read

Original authors: Jiawei Qiu, Yichen Xu, Jianzhe Ma, Mingyang Yu, Wenbin Zhu, Yang Han, Pinzheng Lv, Wenxuan Wang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are watching a movie, but instead of just seeing what happens on screen, you have a super-smart friend sitting next to you who can predict the plot twists before they happen. In the world of artificial intelligence, this friend is called a Multimodal Large Language Model (MLLM). These are powerful computer brains that can "see" images and videos and "read" text, allowing them to understand the world in a way that feels almost human. For a long time, scientists have been teaching these AI friends to spot bad things after they happen, like recognizing a picture of a fire or reading a text about a dangerous animal. But the real magic—and the real safety challenge—lies in proactive thinking. This is the ability to look at a scene, spot a tiny, subtle clue (like a wobbly ladder or a runner leaning too far), and shout, "Wait! Something bad is about to happen!" before the disaster actually strikes. It's the difference between calling the fire department after the house burns down versus noticing the smoke and calling them before the flames start.

This is exactly what a team of researchers from Renmin University of China wanted to test. They asked a simple but scary question: Can our current AI super-brains actually predict physical accidents in real-time, or are they just pretending to be smart? To find out, they built a new testing ground called SPRINT (Sports Proactive Risk INference Testbed). They gathered nearly 3,000 real videos of sports accidents—from skateboarding crashes to gymnastics falls—and paired them with safe videos where nothing bad happened. They didn't just ask the AI, "Did someone get hurt?" (which is easy to answer after the fact). Instead, they asked, "What is about to go wrong, and why?" They wanted to see if the AI could spot the danger while it was still brewing, like noticing a skateboarder's foot slipping before they hit the ground.

The results were a bit of a wake-up call. The researchers found that while these AI models are excellent at spotting danger once it's obvious (like seeing a person lying on the ground), they are surprisingly bad at predicting it before it happens. When the AI was given a video of a sports accident and asked to describe the situation without being told to look for danger, it often missed the warning signs entirely. Even worse, when the researchers explicitly asked, "Is there danger here?", the AI started panicking. It began seeing danger in safe videos, like a gymnast doing a perfect flip or a runner sprinting, simply because the question made it paranoid. It was as if the AI was so eager to say "I found a problem!" that it started inventing problems where none existed.

The study, which analyzed 2,888 videos across 14 different sports, revealed a sharp gap between the AI's ability to sense danger and its ability to understand it. The best AI models could correctly signal that a hazard was present over 95% of the time when the accident was already happening. However, when asked to explain why the accident was happening or to predict it from the very first few seconds of a video, their performance plummeted to below 50%. In fact, when the researchers tested the AI on safe videos using a prompt that asked about danger, some models started flagging harmless high-energy movements as accidents more than half the time.

The researchers also tried to "teach" one of the AI models using their new dataset, hoping to fix these blind spots. After some training, the AI got much better at spotting the causes of accidents and became less likely to cry wolf on safe videos. However, even with this help, the AI still struggled with the hardest part: looking at a video of a sports action in its early stages and correctly identifying the specific physical reason an accident might occur without being prompted to do so.

In the end, the paper suggests that while our current AI friends are getting better at describing what they see, they haven't quite mastered the art of "proactive safety." They are still mostly reactive, waiting for the crash to happen before they understand it. The study concludes that for AI to be truly safe in the real world—whether it's driving a car, helping an elderly person, or watching a sports game—it needs to move beyond just spotting obvious hazards. It needs to learn to understand the causes of danger and to trust its own eyes rather than just reacting to the questions we ask it. Until then, we might need to keep a human eye on the ball, just in case the AI decides a perfect jump is actually a disaster waiting to happen.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →