← Latest papers
🤖 AI

SONIC-O1: A Real-World Benchmark for Evaluating Multimodal Large Language Models on Audio-Video Understanding

The paper introduces SONIC-O1, a comprehensive, human-verified benchmark comprising 60 hours of real-world audio-video data across 13 domains, designed to evaluate and reveal performance disparities in multimodal large language models' capabilities for summarization, question answering, and temporal reasoning.

Original authors: Ahmed Y. Radwan, Christos Emmanouilidis, Hina Tabassum, Deval Pandya, Shaina Raza

Published 2026-05-28
📖 5 min read🧠 Deep dive

Original authors: Ahmed Y. Radwan, Christos Emmanouilidis, Hina Tabassum, Deval Pandya, Shaina Raza

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are hiring a new assistant to help you manage your life. You have two types of candidates: The Super-Expert (a closed-source model like Gemini) and The Open-Source Community Team (various open models like Qwen or MiniCPM).

Until now, you've mostly tested these assistants by showing them static photos (like a family portrait) and asking, "What's happening here?" But real life isn't a photo; it's a movie with sound. It's a conversation where tone of voice, pauses, background noise, and who is speaking when all matter.

The paper introduces SONIC-O1, a new, rigorous "final exam" designed specifically to test how well these AI assistants understand real-world video and audio conversations.

Here is the breakdown of the exam and what the results tell us, using simple analogies:

1. The Exam: SONIC-O1

Think of SONIC-O1 as a 60-hour library of real-life recordings. It's not made-up scripts; it's actual footage from 13 different real-world scenarios, like:

  • The "High-Stakes" Rooms: Courtrooms, emergency response calls, and mental health counseling.
  • The "Daily Grind" Rooms: Job interviews, customer service calls, and doctor visits.
  • The "Community" Rooms: Town hall meetings and sports commentary.

The exam checks three specific skills:

  • The Summarizer: Can the AI watch a 30-minute video and write a clear, accurate summary of what happened? (Like a court stenographer).
  • The Detective (MCQ): Can the AI answer tricky multiple-choice questions that require connecting a visual clue with a spoken clue? (e.g., "Why did the customer get angry?" requires hearing the tone and seeing the gesture).
  • The Time-Tracker (Temporal Localization): This is the hardest part. If you ask, "When exactly did the doctor mention the diagnosis?" the AI must point to the exact second it happened on the clock. It's like asking a referee to say, "The foul happened at 42.3 seconds," not just "sometime in the second half."

Crucially, the exam also checks for fairness. The library includes people of different races, genders, and ages. The researchers want to know: Does the AI work just as well when the speaker is a Black woman in her 20s as when it's a White man in his 40s?

2. The Results: Who Passed?

The Gap Between "Pro" and "Amateur"

  • The Super-Expert (Gemini 3.0 Pro): This model consistently aced the test. It was the only one that could reliably track time in long videos and handle complex, emotional conversations without getting lost.
  • The Open-Source Team: The best open-source models (like Qwen3-Omni) did a decent job on the "Summarizer" and "Detective" tasks. However, they hit a massive wall with the "Time-Tracker" task.
    • The Analogy: Imagine asking an open-source model, "When did the goal happen in the soccer game?" Instead of saying "At 42 minutes," it often says, "At 42 seconds" (forgetting the video started at 0:00 and treating the clip as if it started fresh). They lose their place in the timeline.
    • The Gap: The best open-source model was 22.6% worse than the closed-source model at tracking time.

The "Fairness" Problem
This is where the exam got serious. The researchers found that the AI's performance wasn't the same for everyone.

  • The "Time-Tracker" Bias: The gap in performance was huge depending on who was speaking. For example, when identifying events involving Indigenous speakers, some open-source models dropped to 0% accuracy (they completely failed), while the closed-source model still managed to get about 40% right.
  • The "Black Speaker" Gap: Models consistently performed worse when summarizing or answering questions about videos featuring Black participants compared to Asian or White participants.
  • The Takeaway: The AI isn't just "bad" generally; it has specific blind spots based on who is in the video. It's like a security camera that works perfectly in a well-lit room but goes blind in a corner with a specific type of lighting.

3. The "Audio" Surprise

Many previous tests treated audio as optional, like reading a transcript of a movie instead of watching it.

  • The Finding: When the researchers forced the models to listen to the actual audio (not just text), performance went up.
  • The Catch: The bigger, more powerful models benefited the most from hearing the audio. The smaller models often got confused or didn't improve much. It's like giving a complex musical score to a master musician (they get better) versus a beginner (who might just get overwhelmed).

4. The "Emotion" Check

The researchers also asked the AI to write summaries that sounded empathetic (understanding feelings).

  • The Result: The models had very different "personalities."
    • Gemini sounded like a clinical doctor: serious, factual, but acknowledging emotions.
    • Baichuan and OLA sounded like warm friends: very supportive and positive.
    • Smaller models sounded like robots reading a manual: they could list facts but couldn't really "feel" the emotional tone of the conversation.

Summary

SONIC-O1 is a reality check for the AI world. It proves that while AI is getting good at looking at pictures and reading text, it is still struggling to watch a movie, listen to the sound, keep track of time, and treat everyone fairly.

  • Closed-source models (like Gemini) are currently the only ones reliable enough for high-stakes, real-time understanding.
  • Open-source models are catching up on general understanding but are still "hallucinating" time and failing to be fair to certain demographic groups.
  • Audio matters: You can't just read the subtitles; you have to listen to the voice to truly understand the conversation.

The paper concludes that until we fix these "time-tracking" and "fairness" gaps, we shouldn't fully trust these AI systems to make critical decisions in real-world scenarios like healthcare or law.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →