← Latest papers
💬 NLP

Neural Machine Translation for Coptic-French: Strategies for Low-Resource Ancient Languages

This paper presents the first systematic study on translating Coptic into French, demonstrating that fine-tuning with a stylistically varied and noise-aware corpus significantly improves translation quality and offers valuable insights for low-resource ancient languages.

Original authors: Nasma Chaoui, Richard Khoury

Published 2026-05-14
📖 5 min read🧠 Deep dive

Original authors: Nasma Chaoui, Richard Khoury

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a very old, dusty, and slightly damaged book written in an ancient language called Coptic. You want to translate it into modern French so more people can read it. The problem? There are very few people who speak Coptic today, and there aren't many existing computer programs (AI) that know how to do this translation.

This paper is like a recipe book where the authors, Nasma Chaoui and Richard Khoury, test different "cooking methods" to see how to build the best possible translator for this specific task. They aren't just guessing; they are running experiments to find the most reliable way to turn ancient text into modern French.

Here is a breakdown of their journey, using simple analogies:

1. The Big Question: How do we start?

The authors asked four main questions, which they treated like four different paths up a mountain:

  • Path A: The Direct Route (Fine-Tuning). Take a smart AI that knows a little bit about languages and "train" it specifically on Coptic and French. It's like taking a general student and hiring a private tutor just for Coptic.
  • Path B: The Pivot Route (The Detour). Translate Coptic to English first, then translate that English result into French. It's like trying to get to Paris by first flying to London and then taking a train.
  • Path C: The "Magic Prompt" Route. Ask a smart AI that knows Coptic and English to just "speak French" without any extra training. It's like asking a chef who only knows how to cook Italian to suddenly make a perfect French dish just by asking nicely.
  • Path D: The Multilingual Route. Use a giant AI that knows 100 languages but isn't specialized in any of them. It's like using a Swiss Army knife to try to perform surgery; it has the tools, but maybe not the precision.

The Result: The "Direct Route" (Path A) won hands down. Training a model specifically for Coptic-to-French was much better than using detours or hoping a general AI would figure it out.

2. Choosing the Right "Student" (The Model)

Once they decided to train a model, they had to pick which one to start with. They tested four different "students":

  • The Specialist: An AI that already knew Coptic but only spoke English.
  • The Polyglot: An AI that knew Coptic and French but was part of a huge group of 100 languages.
  • The Generalist: An AI that knew French but didn't know Coptic at all.
  • The Cousin: An AI trained on Hieroglyphs (the ancient ancestor of Coptic).

The Result: The Polyglot (the one that already knew both Coptic and French) was the best student. However, the Cousin (the Hieroglyph expert) did surprisingly well! This is like finding out that if you teach someone to read ancient Egyptian hieroglyphs, they pick up Coptic much faster than someone who has never seen an ancient language before. It proves that knowledge of a "cousin" language helps a lot.

3. The Power of Variety (Multiple Translations)

Ancient texts are tricky because one sentence in Coptic might be translated into French in three different ways, all of which are correct. The authors wondered: Should we teach the AI with just one translation, or all three?

They tested this by training the AI on just one version of the Bible translation, and then on all three versions mixed together.

The Result: The AI trained on all three versions was the most flexible and human-like. It learned that there isn't just one "right" way to say something. It became better at capturing the meaning rather than just copying words. It's like teaching a student by showing them three different ways to write an essay; they learn the core idea better than if you only showed them one rigid template.

4. Making the AI "Tough" (Handling Noise)

Old manuscripts are often damaged. Letters might be missing (holes in the page), smudged, or misread by scanners. The authors wanted to know: Can we make the AI tough enough to handle these mistakes?

They simulated this by intentionally "breaking" their training data. They randomly deleted letters, swapped them, or replaced them with similar-looking ones, just like a damaged manuscript.

The Result:

  • If you train the AI on perfect, clean text, it fails miserably when it sees a damaged page.
  • If you train the AI on 50% damaged text, it becomes a superhero. It learns to guess the missing pieces and ignore the smudges.
  • The sweet spot was training with 50% noise. This created a translator that was robust enough to handle real-world, dusty, damaged ancient books without falling apart.

The Final Takeaway

The authors successfully built the first reliable translator for Coptic to French. Their secret sauce wasn't just one thing; it was a combination of:

  1. Training a model specifically for this pair (not using a detour).
  2. Starting with a model that already knew the ancient language or its "cousin" (Hieroglyphs).
  3. Feeding it many different ways to say the same thing (stylistic variety).
  4. Deliberately "ruining" half of the training data so the AI learned to be tough against damage.

They released these two best models (one based on the Polyglot and one on the Hieroglyph expert) so that Egyptologists and historians can finally use AI to unlock the secrets of ancient Coptic manuscripts, even if those manuscripts are a bit beat up.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →