A Systematic Evaluation of Positional Bias in Multi-Video Summarization with MLLMs
This paper systematically evaluates positional bias in multi-video summarization across nine MLLMs using a new benchmark, revealing that summary quality is significantly influenced by input order and domain, with current mitigation strategies failing to fully eliminate these biases.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a chef preparing a tasting menu for four different dishes: a soup, a salad, a steak, and a dessert. You ask a very smart, but slightly distracted, AI assistant to write a short, perfect description for each dish.
You might think the AI would describe the steak perfectly regardless of whether it was the first or the last dish on the menu. But this paper, titled "A Systematic Evaluation of Positional Bias in Multi-Video Summarization with MLLMs," discovered that the AI is actually quite picky about where a video sits in the lineup.
Here is the breakdown of their findings using simple analogies:
1. The "Seat at the Table" Problem
The researchers found that if you show an AI four videos at once and ask it to summarize each one, the quality of the summary changes depending on the video's position (1st, 2nd, 3rd, or 4th).
- The "Middle-Child" Syndrome: Just like a middle child might get less attention than the firstborn or the youngest, the videos in the middle of the list (positions 2 and 3) often get worse summaries. The AI tends to forget details or write vaguer descriptions for them compared to the videos at the very beginning or the very end.
- The "First Impression" vs. "Last Hurrah": Sometimes the AI loves the first video (the "Primacy" effect), and sometimes it loves the last one (the "Recency" effect). But often, it just ignores the middle ones.
2. The "Magic Trick" of Hiding Bias
The paper points out a tricky part: You can't just look at the average score to see if the AI is biased.
- The Analogy: Imagine a student who gets an A on the first test, an F on the second, an F on the third, and an A on the fourth. Their average grade looks okay (a B), but they clearly have a problem with the middle tests.
- The researchers found that some AI models have a "neutral" average score, but if you look closer, they are actually failing the middle videos while doing great on the ends. They created special tools (metrics) to catch this "Middle-Child Weakness" that standard tests miss.
3. It's Not Just About "More Eyes"
The team tested if giving the AI more resources would fix the problem.
- More Frames: They showed the AI more pictures (frames) from each video.
- More Words: They told the AI it could write longer summaries.
- The Result: It didn't help much. Even when the AI had more "visual budget" or more space to write, it still ignored the middle videos. It's like giving a distracted student a bigger notebook; they still won't write about the middle chapters.
4. The "Order Matters" Experiment
To prove this wasn't a fluke, they used a cyclic rotation method.
- The Analogy: Imagine four friends (Video A, B, C, D) taking turns sitting in four different chairs (1, 2, 3, 4).
- Round 1: A sits in Chair 1, B in 2, C in 3, D in 4.
- Round 2: A sits in Chair 2, B in 3, C in 4, D in 1.
- And so on.
- By doing this, they ensured that every video got to sit in every chair. They found that Video A got a great summary when it sat in Chair 1, but a bad summary when it sat in Chair 3. The video content didn't change, but the seat did.
5. Trying to "Fix" the AI with Prompts
The researchers tried to "coach" the AI to be fair.
- The "Be Fair" Instruction: They added a note to the AI's instructions saying, "Please pay equal attention to every video."
- The Result: It didn't really work. The AI still favored the edges. It's like telling a tired student, "Please study all chapters equally," and they still end up skimming the middle ones.
- One-at-a-Time: They tried asking the AI to summarize only one video at a time, even though it saw all four. This helped a little bit with the middle videos, but it wasn't a perfect fix.
6. The "Leakage" Issue
Finally, they found that sometimes the AI gets confused and mixes up the stories.
- The Analogy: If you ask the AI to describe the "Dessert" video, it might accidentally describe a detail from the "Soup" video because it got mixed up about which video was which. This happens more often when the videos are listed in a long row.
The Bottom Line
The paper concludes that current AI models are not reliable when asked to summarize multiple videos at once. They are sensitive to the order in which you show them the videos. If you want a good summary, you can't just dump four videos in a pile; the position of each video matters a lot, and the AI currently struggles to treat them all equally.
The researchers have built a new "test track" (benchmark) to help future AI developers fix this "positional bias" so that one day, an AI can summarize a playlist of videos without forgetting the ones in the middle.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.