← Latest papers
💬 NLP

CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition

This paper introduces CLBench-V, a comprehensive benchmark designed to evaluate multimodal context learning across three key dimensions—grounding, new information application, and knowledge acquisition—revealing that current state-of-the-art models still struggle significantly with these tasks despite advancements in multimodal capabilities.

Original authors: Lai Wei, Chengqi Li, Jiapeng Li, Ruina Hu, Yue Wang, Weiran Huang

Published 2026-07-29
📖 4 min read☕ Coffee break read

Original authors: Lai Wei, Chengqi Li, Jiapeng Li, Ruina Hu, Yue Wang, Weiran Huang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a super-smart robot to solve a mystery. In the old days, you'd just feed the robot a massive library of books and hope it memorized every fact inside. But in the real world, problems don't come with pre-packaged answers in a library. Instead, you hand the robot a fresh, messy folder containing a map, a financial report, a scientific diagram, and a few photos, and say, "Figure this out using only what's in here." This is the challenge of context learning: the ability to learn new rules or find specific clues from information provided right at that moment, rather than relying on what the robot already knows from its training.

For a long time, scientists tested this skill using only text, like giving the robot a long story to read. But real life is rarely just text. It's a mix of words, charts, maps, and images. The big question researchers are asking is: Can our most advanced AI models actually "read" a multimodal folder (text + images) and learn from it, or do they just pretend to understand while secretly guessing based on old habits? This is the corner of science where artificial intelligence meets the messy, visual reality of the human world.

Enter CLBench-V, a new "stress test" designed by researchers to see how well these AI models handle this specific challenge. Think of CLBench-V not as a single exam, but as a three-level obstacle course.

Level 1: The Detective's Eye (Context Grounding)
First, the model has to simply find the clues. If you show it a subway map and ask, "Which station is between A and B?", it needs to actually look at the map, not just guess based on what it thinks a subway usually looks like. This level tests if the AI can locate and bind the right visual evidence to the question.

Level 2: The Calculator (New Information Application)
Next, the model has to use the clues it found. Imagine the map says, "The ticket costs $5 today," and the question is, "How much for two tickets?" The model needs to take that specific, new number from the image and do the math, ignoring its old memory that tickets might usually cost $3. This tests if the AI can apply fresh facts without getting confused by its prior knowledge.

Level 3: The Scientist (New Knowledge Learning)
Finally, the hardest level. The model must learn a new rule from the context. For example, a scientific paper might show a chart and conclude, "In this specific experiment, blue light makes plants grow faster." The model has to read that, understand the rule, and apply it to a new question about a different plant. It's not just finding a number; it's learning a new law of the universe for the duration of the test.

The researchers built this test by mixing existing public datasets with brand-new tasks they created themselves, specifically focusing on tricky areas like financial reports (calculating Return on Equity), scientific papers (inferring conclusions from medical figures), and sports scenes. They ran 3,443 of these tests on six of the smartest multimodal models available today.

The results? The robots are still struggling. Even the best-performing model, InternVL3.5-30B-A3B, only scored a 0.2847 overall. That's a very low score, suggesting that multimodal context learning is far from being "solved." The paper suggests that while InternVL3.5-30B-A3B is excellent at finding clues (Level 1 / Context Grounding) and learning new rules (Level 3 / New Knowledge Learning), it actually struggles on Level 2 (New Information Application), where models must apply specific new facts from documents. In contrast, the Qwen3.5-Plus model performed best specifically on Level 2 tasks.

Interestingly, the study found that the length of the document or the number of images didn't always make things worse in a simple way. Sometimes, more images actually helped one model but confused another. The biggest issue wasn't just that the documents were long; it was that the models often failed to distinguish between what was in the document and what they already "knew" from their training. They would see a new fact in a chart but then override it with an old guess.

The paper concludes that we are still in the early days. Just because a model can "see" an image and "read" text doesn't mean it can truly learn from them together. The authors hope this new benchmark, CLBench-V, will help developers stop guessing and start fixing these specific blind spots, so our AI can eventually become a true partner in solving complex, real-world problems that require reading the fine print on a map, a chart, or a report.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →