← Latest papers
🤖 AI

KoVRE: Training an Efficient Embedding Model for Korean Visual Document Retrieval

The paper introduces KoVRE, an efficient single-vector embedding model for Korean Visual Document Retrieval that leverages targeted bilingual supervision and advanced training strategies to outperform larger backbones and multi-vector baselines without requiring massive scale.

Original authors: Yongbin Choi, Gyuho Shim, Youngjoon Jang

Published 2026-08-04
📖 3 min read☕ Coffee break read

Original authors: Yongbin Choi, Gyuho Shim, Youngjoon Jang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to find a specific page in a massive, messy library where the books are made of pictures, not just words. If you ask a librarian for "a page with a red chart and a story about money," they might try to read the text out loud first. But what if the text is blurry, or the chart is drawn in a way that words can't describe? That's where Visual Document Retrieval comes in. Instead of just reading the words, this technology looks at the actual image of the page—the layout, the colors, the tables, and the pictures—to find exactly what you need. It's like having a librarian who can "see" the whole page, not just read the letters.

Usually, these smart librarians are trained mostly on English books. If you ask them about Korean documents, they might get confused because the shapes of the letters and the way the pages are designed are different. Also, making these librarians super smart often requires them to be huge, slow, and expensive to run, like a giant supercomputer trying to remember every single pixel of every book. The big question is: Can we build a smaller, faster, and cheaper librarian that is specifically trained to understand Korean visual documents without needing to be a giant?

This is exactly what the researchers behind KOVRE set out to do. They created a new, compact "librarian" (a computer model) specifically for Korean visual documents. Instead of making the librarian bigger, they taught it smarter. They trained it on a massive collection of 708,729 pairs of questions and document pages, mixing both Korean and English to keep the model sharp. They used a clever two-step training process: first, they showed the model tricky examples (hard negatives) to teach it what not to pick, and second, they had a "teacher" model (a very smart but slow expert) show the student model how to rank the best answers.

The results were surprising and promising. Their new 2-billion-parameter model (which is relatively small) didn't just do okay; it actually beat a much larger 8-billion-parameter model and even outperformed a powerful system that uses a complex, memory-heavy method to store information. The paper suggests that by carefully designing the training recipe—specifically how they handled difficult examples and how they normalized scores during the "teacher-student" learning phase—you can create a highly effective Korean document finder without needing a massive computer or a huge storage system. It's a proof that with the right training strategy, a small, efficient model can be just as good, or even better, than the giants.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →