JAMMEval: A Refined Collection of Japanese Benchmarks for Reliable VLM Evaluation
This paper introduces JAMMEval, a refined collection of Japanese benchmarks created through systematic human annotation to address data quality issues in existing datasets, thereby enabling more reliable and accurate evaluation of vision-language models.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to judge how good a new student is at reading a map and finding a hidden treasure. You give them a test. But what if the test itself is broken?
- The Map is Missing: Some questions ask about things that aren't even on the map.
- The Questions are Vague: Some questions are like, "Tell me everything about this picture," which has no single right answer.
- The Answer Key is Wrong: Sometimes the teacher's answer key says "Blue" when the picture clearly shows "Red."
If you use this broken test, you might think the student is bad at maps when they are actually great, or you might think two students are equally good when one is clearly better.
This is exactly the problem the paper "JAMMEval" solves for Japanese AI.
The Problem: The "Broken Test"
Vision-Language Models (VLMs) are AI brains that can "see" pictures and "read" text. To make them smarter, researchers need to test them. However, the existing tests for Japanese AI were like the broken treasure hunt above. They were full of:
- Ambiguous questions (too open-ended).
- Tricks where you could answer without even looking at the picture.
- Wrong answers in the official key.
Because of these flaws, the scores didn't tell the truth. It was like grading a math test where the teacher made a typo in the answer key, punishing the student for being right.
The Solution: JAMMEval (The "Refined Test")
The authors created JAMMEval. Think of this as a team of expert editors who took seven existing, messy test papers and completely cleaned them up.
They didn't just throw away the bad questions (which would have made the test too short). Instead, they fixed them:
- Clarified the Vague: They turned "Tell me about this" into "What color is the bird in the top left corner?"
- Fixed the Answer Keys: They corrected the wrong answers.
- Removed the Cheats: They deleted questions that could be answered without looking at the image.
They did this twice: once by the paper's authors, and a second time by outside experts to make sure everything was perfect.
The Results: A Clearer Picture
After fixing the tests, they ran the same AI models through the new, clean JAMMEval. Here is what happened:
- The Scores Went Up (But for the Right Reasons): The models got higher scores. This wasn't because the models got smarter overnight; it's because the "broken answer keys" were fixed, so the models finally got credit for the right answers they were already giving.
- The Noise Disappeared: Before, if you ran the test twice, the scores would jump around wildly (like a shaky camera). After refining, the scores were steady and reliable (like a steady tripod).
- Better Sorting: The test became better at telling the difference between a "beginner" AI and an "expert" AI. Before, the broken test made them all look the same. Now, the differences are clear.
The "Classroom" Analogy
Imagine a classroom where the teacher (the AI) is being graded by a grader (the benchmark).
- Old Way: The grader is tired, has a bad eye, and sometimes marks "Correct" as "Wrong." The teacher gets frustrated and stops trying to improve because the feedback is confusing.
- JAMMEval Way: The grader gets a new pair of glasses, a fresh answer key, and a clear rubric. Now, when the teacher gets a question right, they get a gold star. When they get it wrong, they get specific feedback on what to fix.
Why This Matters
This paper is a gift to the AI community. By releasing these cleaned-up tests and the code to make them, the authors are saying: "Let's stop arguing about who is the best AI because the test was broken. Let's use a fair test so we can actually see who is improving and who needs more work."
It's a move from "guessing who is good" to "knowing who is good."
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.