← Latest papers
💬 NLP

WAON: Large-Scale Japanese Image-Text Pair Dataset for Improving Model Performance on Japanese Cultural Tasks

This paper introduces WAON, a large-scale Japanese image-text dataset containing approximately 155 million examples, along with a curated benchmark (WAON-Bench), to address the scarcity of high-quality Japanese cultural data and demonstrate that fine-tuning on this dataset significantly improves vision-language model performance on Japanese cultural tasks.

Original authors: Issa Sugiura, Shuhei Kurita, Yusuke Oda, Daisuke Kawahara, Yasuo Okabe, Naoaki Okazaki

Published 2026-04-03
📖 4 min read☕ Coffee break read

Original authors: Issa Sugiura, Shuhei Kurita, Yusuke Oda, Daisuke Kawahara, Yasuo Okabe, Naoaki Okazaki

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a brilliant, world-traveling student how to understand Japanese culture. This student has already read millions of books and seen billions of pictures from around the globe (thanks to a massive dataset called SigLIP). They are smart, but they don't quite "get" the nuances of Japan yet. They might know what a "temple" is, but they might confuse a specific Japanese festival with a generic party, or mistake a local dish for a generic sandwich.

To fix this, you need to give them a specialized textbook filled with Japanese pictures and Japanese descriptions.

This paper introduces WAON, which is exactly that textbook. Here is the breakdown in simple terms:

1. The Problem: The "Lost in Translation" Gap

Currently, if you want to teach an AI about Japan, you have two bad options:

  • Option A: Use a tiny, old library (existing small datasets). It's not enough to teach the AI properly.
  • Option B: Take a giant library of English pictures and use a robot translator to write Japanese captions for them. The problem? The robot translator often makes mistakes, and the pictures themselves are still mostly Western. It's like teaching someone about Tokyo by showing them pictures of New York with Japanese labels stuck on them.

2. The Solution: WAON (The "Native Speaker" Dataset)

The authors built WAON (Web-scale image text Aligned Open Nihongo).

  • Where did it come from? They didn't translate anything. They went straight to the source: the Japanese internet (Common Crawl).
  • What is it? It's a massive collection of 155 million pairs of images and Japanese text.
  • How did they clean it? Imagine sifting through a giant pile of sand to find gold. They used a multi-step filter:
    • They threw out broken links and empty pages.
    • They deleted duplicate photos (like the same logo appearing on 1,000 different websites).
    • They used a "smart eye" (an AI model) to check if the picture actually matched the Japanese text. If the text said "Sushi" but the picture was a "Car," they threw it away.
    • They filtered out bad or unsafe content.

The result is a "pure" dataset where the images and text are naturally Japanese, just like a native speaker would create them.

3. The Test: WAON-Bench (The "Final Exam")

You can't just say, "We made a great textbook!" You have to prove the student learned.

  • The Old Exam: There was an existing test called "Recruit," but it was flawed. It had too many questions about food and not enough about other things. Also, some pictures didn't match the answers (e.g., a picture of a random woman labeled as "Kiyomizu Temple").
  • The New Exam (WAON-Bench): The authors built a brand new, fair test.
    • They created 374 categories (from animals to festivals to buildings).
    • They manually checked every single image to make sure it actually matched the label.
    • It's like replacing a messy, biased quiz with a perfectly curated, comprehensive final exam.

4. The Results: The "Graduation"

They took the smart world-traveling student (the base AI model) and gave them a crash course using the new WAON textbook.

  • The Outcome: The student aced the Japanese cultural test. They scored higher than any other model that had been trained on existing Japanese data.
  • The Surprise: Even though they focused on Japan, the student didn't forget the rest of the world. They actually got better at understanding global things too, proving that learning a specific culture deeply helps you understand the world better.

The Big Picture Analogy

Think of the AI as a chef.

  • Old Method: The chef knows how to cook global cuisine but tries to learn Japanese food by reading English cookbooks that were poorly translated. The dishes come out weird.
  • WAON Method: The chef is given a massive, authentic library of Japanese recipes and photos taken by Japanese grandmothers.
  • WAON-Bench: A panel of Japanese food critics taste the dishes to see if they are authentic.
  • Result: The chef becomes the best Japanese cook in the world, without losing their ability to cook Italian or French food.

In short: The authors built the largest, cleanest, most authentic "Japanese visual dictionary" for AI, and proved that using it makes AI much smarter about Japanese culture than any previous method. They are sharing this dictionary, the test, and the code for everyone to use.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →