Data-Efficient Curation for Multimodal Reasoning under Fixed Training Protocols
This paper investigates data curation strategies for multimodal reasoning under fixed training protocols using the NeurIPS 2025 DCVLR challenge, finding that difficulty filtering on an aligned source corpus yields the most significant accuracy gains—particularly on OlympiadBench—while increasing dataset size primarily reduces variance and other tested diversity or synthetic interventions offer no additional benefit.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a super-smart robot how to solve tricky puzzles that mix pictures and words. You have a very strict rulebook: you cannot change the robot's brain, you cannot change the way you teach it, and you cannot change the tests you give it. The only thing you are allowed to change is the book of practice problems you hand to the robot. This is the world of "data curation" for artificial intelligence. In simple terms, it's the art of picking the right examples to learn from. Most people think that to get smarter, you just need more examples, like reading every book in a library. But what if the library is too big, or you only have time to read a few pages? What if the secret isn't the amount of reading, but picking the perfect few pages that are just hard enough to be interesting, but not so hard that they make the robot give up? This paper asks that exact question: when we are stuck with a fixed teaching method, does the specific choice of practice problems matter more than just grabbing a random pile of them?
The authors of this paper decided to test this idea using a special contest called the DCVLR challenge, which acts like a giant, controlled science lab. They started with a massive collection of practice problems called "Walton," which was already known to be a good fit for the specific robot they were training. Their first big discovery was simple but powerful: if you just grab 1,000 random problems from this collection, the robot gets a decent score. But if you use a "difficulty filter" to pick only the problems that are challenging but still solvable, the robot's score jumps significantly. It's like realizing that for a student preparing for a math test, reading the hardest problems in the book is useless if they can't solve them, and reading the easiest ones is boring because they already know the answers. The sweet spot is the "Goldilocks" zone of problems that make the student think hard but don't make them cry.
The researchers then dug deeper to see if this "difficulty filter" was just a lucky trick for one specific type of test or if it was a real magic bullet. They found that the improvement wasn't just because the filter happened to pick more questions from one popular test (LiveXivTQA). In fact, the biggest boost came from a completely different, massive test called OlympiadBench. This suggests that the strategy of picking "challenging but learnable" problems is a robust way to learn, not just a fluke. They also tried to see if this trick worked on other types of robots (different AI models). It worked for some, but not all, meaning the "difficulty" of a problem depends a little bit on which robot is trying to solve it.
However, the paper also put the brakes on some popular ideas. The team tried adding "diversity" to the mix, thinking that picking problems from many different topics would help. They also tried using "synthetic" problems—questions rewritten by another AI to look longer and more complex. Surprisingly, neither of these tricks helped. In fact, adding these extra layers made things worse or stayed the same. It turns out that in this strict, fixed-recipe setting, trying to be fancy with diversity or rewriting problems didn't add any value. The simple act of filtering for the right level of difficulty was the winner.
Finally, they looked at how many problems the robot needed. They found that once you have about 1,000 of these "just-right" problems, adding more doesn't really make the robot smarter on average. It just makes the robot's performance more stable, so it doesn't have bad days. It's like a chef who has found the perfect recipe with a specific amount of ingredients; adding more of the same ingredients doesn't make the soup taste better, it just ensures the soup tastes the same every time you cook it.
In the end, this paper offers a clear, practical guide for anyone trying to train an AI with limited resources and fixed rules. The recipe is: find a source of problems that matches your robot's style, filter them to find the ones that are tough but doable, and stop there. Don't worry about adding thousands more examples, and don't waste time trying to rewrite the questions or force in too much variety. Sometimes, the best way to learn is simply to pick the right few problems and master them.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.