← Latest papers
💬 NLP

LëtzCross: A Cross-Lingual Page-Level Benchmark for Multimodal Retrieval over Luxembourgish Documents

The paper introduces LëtzCross, a cross-lingual page-level benchmark for Luxembourgish PDFs that demonstrates ColPali-style page-image retrievers outperform OCR-based text-only methods and reveals that multilingual fine-tuning including Luxembourgish significantly boosts retrieval performance.

Original authors: Omar El Bachyr, Fred Philippy, Laura Maria Bernardy, Saad Ezzini, Jacques Klein, Tegawende Bissyande

Published 2026-08-25
📖 5 min read🧠 Deep dive

Original authors: Omar El Bachyr, Fred Philippy, Laura Maria Bernardy, Saad Ezzini, Jacques Klein, Tegawende Bissyande

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the modern digital world, vast libraries of information exist not as neat rows of text, but as complex visual documents: PDFs filled with charts, diagrams, tables, and scattered layouts. For computers to find specific answers within these files, they traditionally relied on a two-step process. First, a system would scan the image of a page and try to read every word, converting the visual ink into digital text. Then, it would search that text for keywords. This method works well for simple documents, but it often stumbles when the answer lies in the shape of a graph or the arrangement of a table, or when the document is written in a language the computer does not know well. This challenge becomes even harder when a user asks a question in one language, like English, but the document they need is written in a rare or low-resource language, such as Luxembourgish. Researchers are now exploring a different approach: teaching computers to look at the document page as a whole picture, understanding the relationship between the words and the visual layout simultaneously, rather than just reading the words.

A team of researchers at the University of Luxembourg set out to test this idea using a specific, difficult case: Luxembourgish. This language, spoken by a small population, has very few digital resources available for computers to learn from. To study how well modern systems could handle this, the team created a new testing ground called L¨etzCross. They gathered hundreds of real Luxembourgish PDF documents, ranging from text-heavy reports to pages rich with visual data like graphs and tables. They then crafted a set of questions that a human might ask about these pages. Some questions required reading the text, while others demanded looking at a chart or a diagram to find the answer. Crucially, these questions were written in four different languages: English, French, German, and Luxembourgish. This setup allowed the researchers to see if a computer could look at a Luxembourgish document and find the right page when asked a question in a completely different language.

The researchers tested two main types of computer systems against this new benchmark. The first group relied on the traditional method: they used software to extract the text from the PDF images and then searched that text. The second group used a newer type of system that looked at the page images directly, treating the visual layout and the text as a single, unified piece of information. When the researchers compared the results, the systems that looked at the images as a whole performed significantly better. They found the correct pages more often than the text-only systems, regardless of which language the question was asked in. The image-based systems were particularly good at answering questions that required understanding visual elements, such as reading a specific value from a bar chart, where the text-only systems often failed because they could not "see" the structure of the data.

To make these systems even better, the researchers experimented with a process called fine-tuning. This is similar to giving a student extra practice tests in a specific subject to sharpen their skills. They trained the image-based systems using questions written in just one language at a time, and also in a mix of all four languages. They discovered that training the system with questions in French actually helped it answer Luxembourgish questions better than training it with German or English questions did. However, the most effective approach was to train the system using a mix of all four languages, including Luxembourgish itself. When the system saw examples of questions in Luxembourgish during its training, its ability to find answers in Luxembourgish documents improved substantially. This suggests that while a system can learn to translate its understanding across languages, seeing the target language directly helps it align its knowledge with the specific documents it needs to search.

The study also looked at how these systems handle the visual part of the document. They tested whether updating the system's ability to "see" the image, in addition to its ability to "read" the text, made a difference. They found that systems which were allowed to adjust both their visual and textual understanding performed better than those that only adjusted their text processing. This confirms that for documents where the layout and graphics carry important information, the computer needs to process the image as a whole, not just the words inside it. While the text-only systems were faster to set up in some cases, the image-based systems provided a more reliable way to retrieve information from complex, visually rich documents, especially when dealing with a language that has limited digital support.

This work provides a clear demonstration that for low-resource languages and complex documents, looking at the picture is often more effective than just reading the text. The researchers found that modern systems capable of understanding both words and images can bridge the gap between different languages more successfully than older methods. By showing that training on a mix of languages, including the target language itself, yields the best results, the study offers a practical path forward for building search tools that work for everyone, not just speakers of major global languages. The findings suggest that as we move toward more visual and multilingual information environments, the future of search lies in systems that can see and understand the document in its entirety.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →