CzechDocs: A Multiway Parallel Dataset of Formatted Documents for Minority Languages in Czechia
This paper introduces CzechDocs, a multiway parallel dataset of formatted documents in Czech and minority languages designed to evaluate and advance machine translation systems capable of preserving document layout.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a beautifully decorated cake. The cake itself is the text you want to translate (the "flavor"), but the frosting, the candles, the little flags, and the fancy piping are the formatting tags (like HTML codes, bold text, or links).
For a long time, when computers tried to translate these cakes, they would often eat the cake, ignore the decorations, and then try to guess where to put the candles back on. Sometimes they put the candles in the frosting, sometimes on the plate, and sometimes they forgot them entirely. This made the final product look messy and broken.
CzechDocs is a new tool created by researchers at Charles University to fix this problem. Here is a simple breakdown of what they did and what they found:
1. The Problem: The "Urgent Cake"
The researchers noticed a real-world need. When thousands of Ukrainian refugees arrived in the Czech Republic in 2022, there was an urgent need to translate government websites, legal forms, and health guides. These aren't just plain text; they are complex documents with buttons, links, and specific layouts. If the translation breaks the layout, the information becomes useless or hard to read.
2. The Solution: A New "Cake Pan" (The Dataset)
The team built a massive collection of 77 unique documents (like government forms and educational guides) that exist in multiple languages at once.
- The Ingredients: They gathered these from real Czech websites.
- The Formats: The documents come in three shapes: HTML (web pages), DOCX (Word files), and PDF (printed brochures).
- The Languages: While the main focus is Czech and Ukrainian, they also included English, Vietnamese, Russian, and others.
- The Magic: They carefully cleaned these documents so that every sentence in Czech has a perfect matching sentence in the other languages, including the exact same formatting tags. It's like having a master blueprint where every candle and frosting swirl is perfectly aligned across all versions.
3. The Experiment: How to Bake the Cake?
The researchers used this dataset to test two different ways of translating these "decorated cakes" using modern AI (Large Language Models):
Method A: The "Strip and Rebuild" (Detag-and-Project)
Imagine taking all the candles and frosting off the cake, giving the plain cake to a baker to translate, and then using a robot arm to try to stick the candles back on exactly where they were before.- Result: This is the traditional way. It works okay, but the robot arm sometimes misses the mark.
Method B: The "Show and Tell" (Direct Tagged Input)
Imagine giving the baker the entire decorated cake and saying, "Translate the cake flavor, but please leave the candles and frosting exactly where they are."- Result: The researchers tested if AI could just look at the whole thing and do it naturally. They found that if you give the AI a specific instruction (a "prompt") telling it to be careful with the decorations, it does a surprisingly good job.
4. The Findings: Who Baked the Best Cake?
They tested two different AI "bakers" (one from OpenAI and one from Cohere) on their dataset.
- The Verdict: Both methods worked well, but they had different strengths.
- The "Strip and Rebuild" method was very consistent at getting the decorations back in the right place, but sometimes the translation of the text itself was slightly less natural.
- The "Show and Tell" method (using the AI's natural ability) produced very high-quality text translations. When the researchers gave the AI a specific instruction to "keep the tags," it did an excellent job of preserving the structure without needing a robot arm to fix it later.
- The Surprise: The AI didn't need to be specially trained to do this; it just needed to be told clearly what to do.
5. What's Next?
The researchers have released the "cleaned" part of their dataset (the validation set) for other scientists to use immediately. They are keeping a "secret test set" (the final exam) for a future competition where different teams will try to see who can translate these decorated documents the best.
In short: They built a high-quality library of real-world, multi-language documents to teach computers how to translate text without breaking the formatting. They found that modern AI is very good at this if you just ask it nicely to keep the decorations intact.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.