Information Router for Mitigating Modality Dominance in Vision-Language Models
This paper introduces \textsc{MoIR} (Multi-modal Information Router), an information-level fusion method that mitigates modality dominance in Vision-Language Models by identifying less informative tokens and routing complementary data from stronger modalities to construct dense representations, thereby improving robustness and performance even under modality degradation.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are hiring a detective to solve a mystery. This detective has two assistants: Vision (who can see photos) and Language (who can read clues).
In the world of modern AI, these detectives are called Vision-Language Models (VLMs). They are incredibly smart, but they have a weird habit: they often ignore the photos and just guess based on the text. This is called "Modality Dominance." It's like the detective saying, "I don't need to look at the crime scene photo; I'll just guess the answer based on the description because reading is easier than looking."
The Problem: The "Lazy" Detective
Previous attempts to fix this were like trying to force the detective to pay attention to the photo. They told the AI, "Hey, look at the picture more!" But here's the catch: sometimes the picture is blurry, dark, or just doesn't have enough clues (low information). If you force the detective to stare at a blank wall, they still can't solve the mystery. They just get confused or make up a story.
The old methods assumed the photo was always full of useful clues. But in the real world, sometimes the photo is just "noise," and the text is the only thing that makes sense.
The Solution: MOIR (The Smart Information Router)
The authors of this paper created a new tool called MOIR (Multi-modal Information Router). Think of MOIR as a super-smart editor or a traffic controller that sits between the assistants and the detective.
Here is how MOIR works, using a simple analogy:
- The Inspection: Before the detective sees the clues, MOIR checks every single piece of information. It asks, "Is this part of the photo actually useful? Or is it just blurry static?"
- The Swap: If MOIR finds a part of the photo that is weak or confusing (like a blurry background), it doesn't just ignore it. Instead, it reaches over to the Language assistant and grabs a strong, clear clue from the text to fill that gap.
- The Result: Now, when the detective looks at the photo, it's no longer a mix of blurry static and clear text. It's a super-charged, information-dense package. Every part of the image has been "boosted" with the best available information from the text.
Why This is a Big Deal
- It Fixes the Root Cause: Instead of just telling the AI "look harder," MOIR actually makes the picture better by filling in the blanks with text.
- It's Robust: Imagine you take a photo of a crime scene, but someone spills coffee on it (corrupts the data). A normal AI would ignore the photo and guess based on the text. MOIR, however, realizes the photo is damaged, grabs the text clues, and "repairs" the mental image so the detective can still solve the case using the actual visual evidence.
- It Prevents Cheating: Without MOIR, the AI might cheat by ignoring the photo entirely. With MOIR, the AI is forced to use the photo because the photo has been made "rich" enough to be useful.
Real-World Results
The researchers tested this on three different types of puzzles:
- Science Questions: Where you need to match a diagram with text.
- Visual Puzzles: Where the photo is messy and hard to understand.
- Video Questions: Where you have to understand a moving scene.
The outcome?
- The AI became much better at solving problems that required looking at the image.
- It stopped "hallucinating" (making up answers) when the image was unclear.
- It became more balanced, using both eyes (vision) and brain (text) equally, rather than relying on just one.
The Bottom Line
Think of MOIR as a translator and a repairman rolled into one. It ensures that when the AI looks at a picture, it's not looking at a blurry mess, but a clear, information-packed version that has been helped out by the text. This makes the AI a much more reliable detective, capable of solving mysteries even when the clues are imperfect.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.