XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models
This paper introduces XModBench, a large-scale tri-modal benchmark designed to evaluate cross-modal consistency in Omni-Language Models, revealing that even state-of-the-art models like Gemini 2.5 Pro struggle with modality-invariant reasoning, exhibit significant performance disparities across modalities, and show systematic directional imbalances.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are teaching a robot to understand the world. You show it a picture of a dog, play a recording of a dog barking, and write the words "dog barking." A truly smart robot should realize that all three things mean the exact same thing, no matter how you present them.
This is the core idea behind XModBench, a new "report card" for the next generation of AI models called Omni-Language Models. These are the super-smart AIs that can see, hear, and read all at once.
Here is a simple breakdown of what the paper is about, using some everyday analogies:
1. The Problem: The "Fashion-Forward" Robot
Current AI models are like a student who is great at reading a textbook but terrible at listening to a lecture or looking at a diagram. They might get the right answer if you ask them a question using text, but if you ask the exact same question using audio or a picture, they might get it wrong.
They aren't actually "understanding" the concept of a dog; they are just memorizing that the word "dog" usually appears next to the answer. They are biased toward the way information is delivered.
2. The Solution: The "Swapping Game" (XModBench)
The researchers created a massive test called XModBench to see if these robots are truly smart or just good at guessing based on the format.
Think of XModBench as a magic mirror game.
- The Setup: They take one single question (e.g., "What animal is making this sound?").
- The Twist: They play the game six different ways:
- Show a picture, ask for the text answer.
- Play a sound, ask for the text answer.
- Show a picture, ask for the sound answer.
- Play a sound, ask for the picture answer.
- And so on...
If the AI is truly smart, it should get the answer right every single time, regardless of whether it's looking, listening, or reading. If it gets the answer right when reading but wrong when listening, the test reveals a "glitch" in its brain.
3. The Test Subjects: 61,000 Questions
The researchers didn't just ask a few questions. They built a giant library of 61,320 questions covering five different areas of life:
- Perception: Recognizing simple things (Is that a cat or a dog?).
- Spatial Reasoning: Understanding where things are (Is the car moving left or right?).
- Temporal Reasoning: Understanding time and order (Did the dog bark before or after the car honked?).
- Language: Translating or understanding emotions in speech.
- External Knowledge: Knowing facts about the world (Who is this singer? What movie is this poster from?).
4. The Results: The "Honest Report Card"
When they ran the top AI models (like Google's Gemini) through this test, the results were eye-opening:
- The Good News: The AIs are getting really good at recognizing things when you show them pictures or text. They are like great readers.
- The Bad News: They are terrible at listening. When the question came as audio, their performance dropped significantly. It's like a student who aces a written test but fails when the teacher asks them to solve a problem out loud.
- The "Directional" Flaw: The AIs also have a weird bias in how they process information. They are great at looking at a picture and guessing the text, but much worse at reading the text and guessing the picture. It's like they are "right-handed" with their brains and struggle when they have to use their "left hand."
5. Why This Matters
This paper is important because it stops us from just looking at the "average score" of an AI. Before, if an AI got 80% right, we thought it was smart. Now, XModBench shows us that maybe it got 95% right on pictures but only 40% right on sounds.
The Big Takeaway:
We are building "Omni" models (models that can do everything), but right now, they are more like "Uni" models (models that are good at only one or two things). XModBench is the tool we need to fix this, forcing developers to teach their AIs to be truly consistent, so they can understand the world whether you show them a photo, play a song, or write a note.
In short: XModBench is the ultimate "lie detector" for AI, ensuring they aren't just memorizing patterns but actually understanding the world behind the pixels, the waves, and the words.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.