CURV: Enhancing Chart Understanding Through Curriculum Visual Grounded Reasoning
The paper introduces CURV, a curriculum learning framework paired with the CCQA dataset, which enhances chart question answering by training multimodal models to perform intrinsic, multi-step visual grounded reasoning, resulting in significant performance improvements across various benchmarks.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to solve a complex puzzle, but instead of just looking at the picture on the box, you have to figure out the picture while you are building it. This is the daily struggle of Multimodal Large Language Models (MLLMs)—super-smart computer programs that can read text and look at images at the same time. Right now, these programs are like brilliant students who can read a math textbook perfectly but get confused when asked to solve a problem on a whiteboard because they can't quite focus on the right numbers. They often mix up what they see with what they think, leading to answers that sound logical but are totally wrong because they misread a chart or missed a tiny detail. Scientists have tried to fix this by giving the computers "reference materials" or asking them to talk through their steps out loud (a method called Chain-of-Thought), but the computers still struggle to connect their thinking directly to the visual evidence in real-time. They need to learn how to look, think, and look again, all in one smooth motion, just like a human does.
Enter CURV, a new training method designed to teach these AI models how to become better visual detectives. Think of CURV as a "video game level system" for learning. Instead of throwing a computer into the hardest boss battle immediately, CURV starts it on "Level 1," where it learns to read simple charts and point to the right numbers. As it gets better, it unlocks "Level 2," where it has to combine two pieces of information, and finally "Level 3," where it has to juggle multiple charts at once. The secret sauce is that at every single step of the way, the model is forced to physically point to the part of the image it is talking about. It's like teaching a child to do math by making them touch the apples on the table before they can say "five."
The researchers behind this paper, CURV, found that this step-by-step, "look-then-think" training works wonders. They created a massive new dataset called CCQA (Curriculum Chart Question Answering) to act as the training ground. This dataset is built like a ladder, with questions getting progressively harder, from simple "what is this number?" to complex "compare these two graphs and tell me the difference." By training on this ladder, the models learned to internalize the skill of grounding their reasoning in what they actually see. The results were impressive: models trained with CURV improved their accuracy by up to 20.50% on their training tests. Even more importantly, they didn't just get good at the practice questions; they got better at real-world charts they had never seen before, improving by up to 12.30%, and even got better at other types of visual puzzles, like math problems, by up to 10.20%.
The paper suggests that the key to this success isn't just giving the AI more data, but changing how it learns. The researchers argue that previous methods failed because they treated seeing and thinking as separate tasks. CURV proves that by weaving them together—making the model constantly shift its visual focus to match its logical steps—it builds a stronger, more reliable brain for understanding charts. However, the authors also note a funny quirk: while forcing the model to "point" at the image during training makes it smarter, forcing it to "point" out loud during the final test sometimes makes it stumble, likely because it gets confused by its own pointing. This suggests that the best approach is to teach the model to look and think together until it becomes a natural habit, so it doesn't need to stop and point anymore to get the right answer. Ultimately, CURV shows that if you teach AI to break big problems into small, visual steps, it can learn to understand the world's data much more like a human does.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.