From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models
This paper presents the first systematic survey of unified vision-language perception in Multimodal Large Language Models (MLLMs), formalizing perception as an intrinsic capability, proposing a five-stage taxonomy of its paradigm evolution, and outlining future research directions toward artificial general intelligence.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are teaching a brilliant but slightly clumsy student how to understand the world. This student is a Multimodal Large Language Model (MLLM). It's great at reading books (text) and looking at pictures (vision), but for a long time, it struggled to connect the two. It could describe a picture vaguely, but it couldn't point to exactly where a cat was sitting or explain why the cat looked scared.
This paper, titled "From Structure to Synergy," is like a history book and a roadmap for how we taught this student to truly "see" and "understand" at the same time. The authors argue that we shouldn't treat vision and language as two separate subjects; instead, they need to be fused into one super-skill called Perception.
Here is the story of how we got there, broken down into five chapters, using simple analogies:
The Problem: The "Blind" Student
Before this survey, most reviews looked at "vision" and "language" separately. It was like having one teacher for math and another for art, but never asking them to work together. The authors say this is a mistake. To truly understand an image, you need to look at it and talk about it simultaneously. They wanted to create a single map of how these models evolved to do exactly that.
The Five Stages of Evolution
The authors divide the history of these models into five stages, moving from "fixing the hardware" to "teaching the brain to think."
Stage 1: Fixing the Eyes (Encoder-Centric)
The Analogy: Imagine the model's "eyes" (the part that looks at the image) were blurry.
What happened: Researchers started by upgrading the eyes themselves. They added special lenses (modules) that could zoom in on specific parts of a picture, like a magnifying glass, before the model even tried to read it. They also tried giving the model multiple pairs of eyes (different experts) to look at the image from different angles.
The Result: The model stopped just seeing the whole picture as a blur and started noticing specific regions, like "that red ball over there."
Stage 2: Fixing the Hands (Decoder-Centric)
The Analogy: Now the eyes could see the regions, but the model's "hands" (the part that draws or points) were clumsy. It could say "the ball is there," but it couldn't draw a perfect circle around it.
What happened: Researchers added special tools to the model's output side. They built "pixel-perfect" hands that could draw exact outlines, masks, or boxes around objects.
The Result: The model went from saying "there is a cat" to being able to draw a perfect outline around the cat's fur, pixel by pixel.
Stage 3: Learning to Look Twice (Dynamic Perception)
The Analogy: Imagine you are looking at a complex painting. You don't just glance once; you zoom in, step back, look at the details, and maybe ask for a magnifying glass.
What happened: The old models looked at an image once and gave an answer. The new models learned to think dynamically. They learned to:
- Call in outside tools (like an OCR scanner to read text in the image).
- Zoom in on specific spots if they were confused.
- Write little computer programs to help them solve the puzzle.
The Result: The model became an active investigator. Instead of a static snapshot, it could "walk around" the image, gathering clues step-by-step.
Stage 4A: Training Without Surgery (Architecture-Free Strategies)
The Analogy: Instead of giving the student new glasses or new hands, we just changed how we taught them.
What happened: Researchers stopped changing the model's physical structure. Instead, they used Instruction Tuning (giving better practice tests) and Reinforcement Learning (like a video game where the model gets points for correct answers and loses points for mistakes).
The Result: The model got smarter at finding objects and solving puzzles just by practicing harder and learning from its mistakes, without needing a hardware upgrade.
Stage 4B: The Grand Synergy (Towards Unified Perception)
The Analogy: This is the "Superhero" phase. The student now combines everything: they have great eyes, great hands, they know how to zoom in, they can write code, and they know how to learn from rewards.
What happened: The paper looks at the newest models (like OpenAI's o3) that blend all these skills. They don't just follow a rigid checklist; they plan, they choose their own tools, they zoom in when needed, and they reason through the problem like a human.
The Result: We are moving toward a model that doesn't just "process" an image but truly perceives it in a unified, flexible way.
The Roadblocks Ahead
Even though we've made huge progress, the authors point out three big hurdles:
- The Data Diet: These models need massive amounts of high-quality, carefully labeled data to learn. It's expensive and hard to get.
- The Grading Problem: It's easy to grade a math test (right or wrong), but it's hard to grade "perception." How do you give a precise score for how well a model "saw" a complex scene? We need better ways to reward the model.
- The Cost of Thinking: The smartest models that "zoom in" and "think twice" require a lot of computer power. It's like running a supercomputer just to look at a photo, which is too expensive for everyday use right now.
The Bottom Line
This paper is a guidebook. It tells us that to build truly intelligent machines, we can't just patch up the eyes or the hands separately. We need to teach the whole system to synergize—to see, think, and act as one unified, adaptive brain. The journey has gone from fixing the hardware to teaching the brain to think, and the next step is making that thinking efficient and truly general.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.