← Latest papers
💬 NLP

Targum - A Multilingual New Testament Translation Corpus

The paper introduces "Targum," a multilingual corpus of 651 New Testament translations across five European languages that offers unprecedented depth and standardized metadata to enable flexible, multi-level quantitative analysis of translation history.

Original authors: Maciej Rapacz, Aleksander Smywiński-Pohl

Published 2026-03-17
📖 5 min read🧠 Deep dive

Original authors: Maciej Rapacz, Aleksander Smywiński-Pohl

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a massive library of books, but instead of just one copy of Harry Potter, you have 651 different versions of it. Some are written in perfect, old-fashioned English; others are written in slang; some are translated from the original French, others from a German version of the French.

For a long time, computer scientists studying language mostly cared about having many different languages in their library (breadth), even if they only had one or two books per language. They treated the Bible like a giant puzzle where every sentence in English matched a sentence in French, just to teach computers how to translate.

But the authors of this paper, Maciej Rapacz and Aleksander Smywiński-Pohl, say, "Wait a minute. We're missing the real story." They realized that to understand how language and culture change, you don't just need many languages; you need deep dives into a few specific ones.

So, they built Targum.

What is Targum?

Think of Targum as a giant, organized museum of New Testament translations.

  • The Collection: They gathered 651 different versions of the New Testament from five major European languages: English, French, Italian, Polish, and Spanish.
  • The Depth: Before this, if you wanted to study English Bible translations, you might have found 40 versions. Targum found 390. That's like finding 10 times more paintings by the same artist to study how their style changed over 400 years.
  • The Name: They named it "Targum" after ancient Aramaic translations that were famous for being both literal and creative—just like the mix of strict and loose translations they collected.

The Problem They Solved: The "Duplicate" Mess

Imagine you go to ten different websites to download a PDF of the same book. You end up with ten files.

  • Website A calls it "The King James Bible (1769)."
  • Website B calls it "KJV 1769 Edition."
  • Website C calls it "King James Version, Revised 1769."

To a computer, these look like three totally different books. To a human, they are the same book.

The authors spent a lot of time acting like detectives. They manually checked publication records, publisher catalogs, and old prefaces to figure out which files were actually the same "soul" (the same translation work) and which were just different copies. They created a "fingerprint" (a unique ID) for every distinct version. This allows researchers to say, "I want to study only the 1769 KJV," or "I want to study every version that came after the KJV."

Why Does This Matter? (The "Why Should I Care?" Part)

1. It's a Time Machine for Language
Because they have so many versions from different years, researchers can watch language evolve.

  • Analogy: Imagine watching a time-lapse video of a forest growing. You can see how the trees (words) changed shape, how the underbrush (grammar) thickened, and how the seasons (historical eras) affected the growth.
  • Real finding: They discovered that in the last 50 years, translations have become much more diverse. Old translations tended to sound very similar to each other, but modern ones are all over the map—some very literal, some very loose.

2. It's a Style Guide for Translators
If a new translator (human or AI) makes a new version of the Bible today, they can use Targum to see where their work fits.

  • Analogy: It's like a musician releasing a new song and using a database to see: "Does my song sound more like a 1960s Beatles track or a 2020s Hip-Hop beat?"
  • Real finding: They found that French translations tend to be longer and more wordy than Polish ones, not because of the Bible, but because of how the French and Polish languages naturally work (like how French uses more small words like "the" and "of").

3. It Helps Fix AI
Big Language Models (like the one you are talking to right now) are trained on huge amounts of text. But they often struggle to understand nuance or history. Targum gives them a specialized dataset to learn how the same idea can be expressed in dozens of different ways, helping AI understand the "flavor" of language, not just the dictionary definition.

The "Fine Print" (Limitations)

The authors are honest about the flaws:

  • The Digital Bias: They only grabbed books that were already online. If a rare 18th-century translation exists only in a dusty basement in a library and hasn't been scanned, Targum doesn't have it. It's a map of the digital world, not the entire world.
  • The Detective Work: Figuring out exactly when a modern, internet-only translation was updated is hard. Sometimes a website updates a file without telling anyone. The authors did their best to guess the dates, but it's not always perfect.

The Bottom Line

Targum is a massive, high-definition map of how people have tried to tell the same story in different ways over centuries.

Instead of just asking, "How do I translate this word from English to French?", this corpus lets us ask, "How has the entire history of English translation changed? How do different cultures handle the same story? And how can we use that history to build better tools for the future?"

It turns the Bible from a simple dataset for training robots into a rich, complex laboratory for understanding human culture.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →