← Latest papers
💻 computer science

Jagle: Building a Large-Scale Japanese Multimodal Post-Training Dataset for Vision-Language Models

This paper introduces Jagle, the largest Japanese multimodal post-training dataset to date comprising 9.2 million instances generated from diverse sources, which significantly enhances the performance of Japanese vision-language models while maintaining or improving their English capabilities.

Original authors: Issa Sugiura, Keito Sasagawa, Keisuke Nakao, Koki Maeda, Ziqi Yin, Zhishen Yang, Shuhei Kurita, Yusuke Oda, Ryoko Tokuhisa, Daisuke Kawahara, Naoaki Okazaki

Published 2026-04-03
📖 4 min read☕ Coffee break read

Original authors: Issa Sugiura, Keito Sasagawa, Keisuke Nakao, Koki Maeda, Ziqi Yin, Zhishen Yang, Shuhei Kurita, Yusuke Oda, Ryoko Tokuhisa, Daisuke Kawahara, Naoaki Okazaki

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot how to understand the world. You give it eyes (a camera) and a brain (a language model). To make this robot smart, you need to show it millions of pictures and tell it stories about what's in them. This is how "Vision-Language Models" (VLMs) are trained.

For a long time, scientists have had a massive library of these picture-stories in English. They could just mix and match thousands of existing books to create a super-smart English robot.

But for Japanese, the library was almost empty. There weren't enough books, and the ones that existed were too small or only covered simple topics. You couldn't just translate the English books because the pictures in them often had English text on them (like signs or charts), which would confuse a Japanese robot.

Enter "Jagle."

Think of Jagle as a massive, custom-built construction project to fill that empty Japanese library. Here is how the authors built it, using some simple analogies:

1. The Problem: The "Empty Shelf"

Most researchers build datasets by gathering existing "Question and Answer" books. In English, there are millions of these. In Japanese, the shelves are bare. Trying to build a smart Japanese robot with so little data is like trying to teach a child to read using only a single comic book.

2. The Solution: Building from Scratch

Instead of looking for existing books, the Jagle team decided to build the books themselves. They gathered raw materials from everywhere:

  • Photos taken in Japan: Like a digital photo album of everyday life.
  • PDFs and Documents: Scanned government papers, textbooks, and reports.
  • Web pages: Billions of image-text pairs from the internet.

They didn't just copy-paste; they used a "factory" approach to turn these raw materials into learning exercises.

3. The Factory: Four Ways to Make Questions

The team used four different "machines" to turn raw data into training questions:

  • The AI Interviewer (VLM-based Generation): They used a super-smart AI (Qwen3-VL) to look at a picture and ask, "What is happening here?" and then answer it. It's like having a robot tour guide who looks at a photo of a police station and says, "This is the Gunma Prefectural Police Headquarters."
  • The Translator: They took existing English charts and graphs, translated the text inside the image and the questions into Japanese, ensuring the picture and the words matched perfectly.
  • The Printer (Text Rendering): For tasks where the robot needs to read text from an image (like reading a sign), they took text, printed it onto a blank image, and asked the robot to read it back.
  • The Scanner (OCR): They used specialized tools to pull text out of messy documents and create questions based on that text.

4. The Result: A Massive Library

The result is Jagle, a dataset with 9.2 million examples.

  • It covers 5 different types of thinking: General questions (What's in this photo?), reading charts, writing descriptions, reading text from images, and simple text extraction.
  • It's the largest Japanese dataset of its kind ever made.

5. The Test: Does it Work?

The researchers trained a small robot brain (2.2 billion parameters) using this new library.

  • Japanese Skills: The robot became incredibly smart at Japanese tasks. It beat the previous best Japanese robot (InternVL3.5) and came very close to the world's best multilingual robot (Qwen3-VL).
  • English Skills: Here is the magic trick. Usually, when you teach a robot a new language, it gets worse at its old language (like a student focusing so hard on French they forget their English). But with Jagle, the robot got better at English too!
    • Why? The authors think that by adding so much diverse Japanese data, the robot learned to be more flexible and creative, which actually helped it solve English problems better than before.

The Big Picture

The paper shows that you don't need to wait for someone else to build a library for your language. If you have the right tools and a creative pipeline, you can build a massive, high-quality training dataset from scratch.

Jagle is like handing a Japanese student a brand-new, 9-million-page encyclopedia of pictures and stories, finally giving them the same chance to become a genius as their English-speaking counterparts. And the best part? They released the library, the robot, and the blueprints for everyone to use.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →