Pre-Editorial Normalization for Automatically Transcribed Medieval Manuscripts in Old French and Latin
This paper introduces the task of Pre-Editorial Normalization (PEN) to bridge the gap between palaeographically accurate and normalized digital editions of Old French and Latin manuscripts, presenting a new dataset and a ByT5-based model that achieves state-of-the-art performance in converting raw Automatic Text Recognition outputs into usable, normalized text.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a time machine that can take a photo of a handwritten page from a medieval manuscript (like a book from 800 years ago) and turn it into digital text. This is what Automatic Text Recognition (ATR) does. It's amazing, but it's like having a very literal, slightly confused robot assistant.
Here is the problem: The robot assistant sees the page exactly as it is written.
- It sees a scribbled "m" and writes "m".
- It sees a tiny squiggle meaning "and" and writes "and" (or sometimes just a weird symbol).
- It sees a word split across a line break and writes it with a hyphen in the middle.
- It sees a fancy "u" that looks like a "v" and writes "v" because that's what the ink looks like, even if it should be a "u".
If you try to use this raw text with modern computer tools (like a spell-checker, a translator, or a search engine), the tools get confused and crash. They expect clean, modern English or French.
On the other hand, if you ask a human scholar to fix it all at once, they might "over-correct." They might change a word just because they think it should be spelled a certain way, accidentally inventing words that were never there (hallucinations) or erasing historical quirks that are actually important.
The Solution: The "Pre-Editorial Normalization" (PEN) Sandwich
The authors of this paper propose a new middle-ground step called Pre-Editorial Normalization (PEN). Think of it as a specialized "translator" or "clean-up crew" that sits between the raw robot scan and the final polished book.
Here is how the process works, using a Restaurant Analogy:
The Raw Ingredients (Graphemic ATR):
Imagine the robot scanner is a chef who just dumped a pile of raw, unpeeled, muddy vegetables onto the counter. Some are whole, some are chopped weirdly, and some have dirt (recognition errors) on them. This is the "graphemic" output. It's accurate to the image, but it's not ready to eat.The Old Way (The "All-in-One" Chef):
Previously, scholars tried to hire one chef to wash, peel, chop, cook, and plate the food all at once. The problem? This chef often got tired, guessed wrong about what the vegetable was, and sometimes added ingredients that weren't in the original recipe (hallucinations).The New Way (The PEN Sandwich):
The authors introduce a new step: The Prep Station.- Step 1: The robot scanner dumps the raw veggies (the manuscript text).
- Step 2 (The PEN Model): A specialized AI (trained on millions of examples) comes in. It doesn't try to cook the meal yet. It just does the prep work:
- It washes off the mud (fixes robot reading errors).
- It peels the skins (expands abbreviations like turning "q̃" into "qui").
- It separates the mixed-up piles (fixes spacing).
- It standardizes the shapes (changes "u" to "v" where it's a consonant, but keeps the original spelling of words like "philosophia" if that's how it was written).
- Step 3: Now, the text is clean and normalized. It's ready for the final "plating" (editing) or for modern computers to analyze.
Why is this a big deal?
The authors built a massive training dataset (a library of 4.6 million pairs of "dirty" robot text and "clean" human text) to teach their AI how to do this prep work.
- The Result: Their AI model is like a master prep chef. It cleaned up the text with 93% accuracy (only 6.7% errors), which is much better than previous attempts.
- The Benefit: Because this AI separates the "cleaning" from the "cooking," we can now trust the computer analysis more. If the computer makes a mistake later, we know it's because of the recipe (the analysis), not because the ingredients were muddy.
The Catch (The "Over-Correction" Risk)
The paper also warns that even this new chef has a bias.
- For Latin: The chef tends to make everything look like "Classical Latin" (the fancy, perfect version from ancient Rome), sometimes erasing the unique, messy medieval spelling that historians actually want to study.
- For Old French: The chef tends to pick the most common way to spell a word, ignoring rare, local dialects.
The Bottom Line
This paper is about building a smart, middle-layer filter for historical texts. Instead of trying to turn a muddy, ancient manuscript directly into a modern book in one giant leap, they break it down. They first turn the muddy text into "clean, standardized text" using a specialized AI. This makes it possible for computers to read, search, and analyze medieval history without getting lost in the mud, while still keeping enough of the original flavor to be historically accurate.
It's the difference between trying to read a handwritten note through a foggy window versus cleaning the glass first so you can see the words clearly.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.