← Latest papers
⚡ electrical engineering

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs

This paper introduces AVI-Bench, a cognitively inspired benchmark featuring a three-stage evaluation framework and a primitive sensation extension (AVI-Bench-PriSe) to systematically assess and diagnose the current limitations in audio-visual intelligence of Omni-MLLMs, ultimately proposing a four-level taxonomy to guide future model development.

Original authors: Yaoting Wang, Ziyi Zhang, Wenming Tu, Shaoxuan Xu, Wenjie Du, Cheng Liang, Weijun Wang, Yuanchao Li, Guangyao Li, Hao Fei, Yuanchun Li, Henghui Ding, Yunxin Liu

Published 2026-06-09
📖 5 min read🧠 Deep dive

Original authors: Yaoting Wang, Ziyi Zhang, Wenming Tu, Shaoxuan Xu, Wenjie Du, Cheng Liang, Weijun Wang, Yuanchao Li, Guangyao Li, Hao Fei, Yuanchun Li, Henghui Ding, Yunxin Liu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot to be a "human-like" observer of the world. You want it to not just see a video or hear a sound, but to truly understand how the two work together—like realizing that the sound of a siren matches the flashing lights of a police car, or that a person's angry voice matches their furrowed brow.

The paper "AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs" introduces a new, rigorous test called AVI-Bench to see how good current AI models really are at this specific skill.

Here is the breakdown of what they did, using simple analogies:

1. The Problem: The "One-Track Mind" Robot

Current AI models (called Omni-MLLMs) are like students who are great at reading textbooks but terrible at listening to music or watching a movie. They can process text, images, and audio separately, but they struggle to mix them together seamlessly.

  • The Gap: Existing tests were like giving the student a math quiz or a music quiz, but never asking them to solve a math problem while listening to a song. We didn't know if they could handle the real world where sight and sound happen at the same time.

2. The Solution: The "Cognitive Gym" (AVI-Bench)

The authors built a new gym for AI models to train and test their "Audio-Visual Intelligence" (AVI). Instead of just checking if the AI gets the right answer, they check how the AI thinks, breaking it down into four levels of difficulty, much like a video game with increasing levels:

  • Level 1: Perception (The Eyes and Ears)
    • The Analogy: Can the AI spot the objects? Can it hear the sounds?
    • The Test: "How many dogs are in this video?" or "Is this sound a car horn or a bird?" It checks if the AI can simply detect what is there.
  • Level 2: Understanding (The Brain's Filing Cabinet)
    • The Analogy: Can the AI connect the dots?
    • The Test: "Describe the scene." or "Find the video clip that matches this audio." It checks if the AI understands the story or the relationship between what it sees and hears.
  • Level 3: Reasoning (The Detective)
    • The Analogy: Can the AI solve a mystery?
    • The Test: "Why is the person running?" or "If the glass breaks, what sound should we hear?" It checks if the AI can make logical conclusions based on the combined clues.
  • Level 4: Primitive Sensation (The "Alien" Test)
    • The Analogy: This is the paper's unique twist. Imagine showing the AI a video of a weird, abstract shape moving with a strange, low-pitched hum. It has no "meaning" (no cars, no people, no words).
    • The Test: "Is the shape getting bigger?" or "Is the sound getting louder?"
    • Why it matters: Most AI is trained on "familiar" data (cats, dogs, cars). This test sees if the AI can react to raw sensory data the way a human baby does, without needing to recognize a specific object first.

3. The Results: The "Visual Specialist" Trap

The authors tested 28 different AI models (including big names like GPT-4o and Gemini). The results were revealing:

  • The "Visual Specialist" Syndrome: Most models are like a person with 20/20 vision but who is hard of hearing. They are great at tasks involving images but struggle significantly when audio is the main clue.
  • The Bottleneck Effect: The paper found that if an AI is bad at seeing or hearing (Perception), it cannot be good at solving mysteries (Reasoning). You can't build a strong reasoning tower on a weak foundation.
  • The "Alien" Failure: When tested on the "Primitive Sensation" (unfamiliar, low-meaning data), the models crashed. They performed terribly compared to humans. It's like giving a human a test on a language they've never heard; they can't guess the rules just by looking at the shapes.
  • Grounding is Hard: Even the best models struggled to point to exactly where a sound is coming from in a video (like pointing to the speaker in a room).

4. The New Scorecard: The Four-Level Taxonomy

Instead of just giving one average score (which can hide weaknesses), the authors created a Four-Level Taxonomy to grade models more fairly:

  1. Task-Adaptive: Can it do the job at all?
  2. Modality-Adaptive: Is it balanced, or is it just a "visual specialist"?
  3. Stage-Adaptive: Does its reasoning actually rely on good perception, or is it just guessing?
  4. Domain-Adaptive: Can it handle weird, unfamiliar situations, or does it only work on things it has seen before?

The Bottom Line

The paper concludes that while AI is getting smarter, it is still far from "human-like" audio-visual intelligence. It is currently a collection of specialists that haven't learned to work together well, especially when faced with strange or new situations.

AVI-Bench is the new ruler they are handing to researchers to measure progress, ensuring that future AI doesn't just memorize answers but actually learns to perceive, understand, and reason like a human does.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →