Responsibility Distribution Estimation in Ego-View Accident Videos with Multimodal Large Language Models
This paper introduces the novel task of responsibility distribution estimation in ego-view traffic accident videos, demonstrating that fine-tuned multimodal large language models can effectively predict agent responsibility percentages by leveraging the driver's direct visual perspective.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to figure out who is to blame for a car accident. Usually, police or insurance companies look at the scene from a "bird's-eye view"—like a security camera on a street corner or a satellite photo. They see the whole picture, but they can't see exactly what the driver inside the car saw right before the crash.
This paper introduces a new way to look at accidents: from the driver's seat.
Here is a simple breakdown of what the researchers did, using some everyday analogies:
1. The Problem: The "Bird's-Eye" Blind Spot
Think of traditional accident analysis like watching a play from the back of the theater. You see the actors (cars and pedestrians) and the stage, but you don't know what the main character (the driver) saw through their specific pair of glasses.
- The Issue: Security cameras are expensive to install everywhere, and they can't tell us if the driver actually had time to react or if something popped up suddenly in their view.
- The Solution: The researchers used "ego-view" videos. This is footage recorded by a dashcam inside the car. It's like putting a camera on the driver's forehead. It shows exactly what the driver saw, making it much fairer to judge if they could have avoided the crash.
2. The New Game: Splitting the Pie, Not Picking a Winner
In most accident studies, the goal is to say, "The pedestrian is 100% wrong" or "The car is 100% wrong."
- The New Task: The researchers created a game called "Responsibility Distribution." Instead of picking one winner, the model has to split a 100% pie.
- Example: Maybe the pedestrian ran out too fast (60% blame), but the driver was also texting and didn't brake hard enough (40% blame). The goal is to teach the AI to slice that pie correctly.
3. How They Taught the AI (The "Teacher" and the "Student")
They didn't have human lawyers label thousands of videos (that would take forever). Instead, they used a clever two-step process:
- Step 1 (The Teacher): They used a very smart AI (a Large Language Model) to watch the dashcam videos and write down a "best guess" for how to split the blame pie. They made sure the numbers always added up to 100%.
- Step 2 (The Student): They took a different AI model (Qwen3-VL) and "fine-tuned" it using those teacher's notes. Think of this as a student studying a teacher's answer key to learn how to grade a test.
4. The Experiment: What Did the AI Need to See?
They tested the AI in three different ways to see what information was most important:
- The Eyes (Video Only): The AI watched the dashcam frames.
- The Glasses (Video + Text): The AI watched the video and read a written description of the scene (like "sunny day," "urban street").
- The Blindfold (Text Only): The AI only read the description and saw no video.
The Results:
- The Winner: The AI that saw the video and read the text was the best. It got the blame split right about 79% of the time (Exact Match).
- The Loser: When they took away the video and only gave the text, the AI got terrible at guessing the blame split. This proves that seeing the video is crucial. You can't understand a driver's reaction just by reading a report; you have to see what they saw.
- The "Glasses" Test: They also tried giving the AI "segmentation overlays" (images where cars and people are colored in solid blocks). This didn't help much more than just showing the raw video, suggesting the AI is smart enough to understand the raw picture on its own.
5. The Big Warning (Ethical Reality Check)
The authors are very clear: This is a research prototype, not a judge.
- The "Teacher" wasn't perfect: The "answer key" used to train the AI was generated by another AI, not a human lawyer. It's a good guess, but not legal truth.
- Missing Details: The AI can't see the car's speedometer or know the specific traffic laws of that country. It only sees what's in the video.
- The Rule: This tool should never be used to automatically decide who goes to jail or who pays insurance. It's like a "second opinion" tool to help human investigators, not a replacement for them.
Summary
The paper shows that if you give a smart AI a dashcam video and ask it, "How much blame should go to the driver vs. the other person?", it can actually do a pretty good job—but only if it can see the video. It's a big step toward understanding accidents from the driver's perspective, but it's still just a tool for research, not a courtroom verdict.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.