← Latest papers
📄 other

Cross-Lingual Transfer Learning and Autonomous Data Bootstrapping for VLM-Based Ottoman Turkish Handwritten Text Recognition

This paper introduces Azra, a Vision-Language Model framework for Ottoman Turkish handwritten text recognition that achieves state-of-the-art performance through LoRA fine-tuning on curated data and demonstrates that autonomous data bootstrapping can effectively match curated approaches while significantly reducing annotation costs.

Original authors: Gökhan Usta, Oğuz Alpoğlu, Fatih Günaydın

Published 2026-07-24
📖 5 min read🧠 Deep dive

Original authors: Gökhan Usta, Oğuz Alpoğlu, Fatih Günaydın

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine trying to read a diary written by your great-great-grandparent, but the handwriting is a tangled, cursive mess, the language has changed completely, and the alphabet looks like a secret code. This is the daily reality for historians trying to unlock the millions of documents from the Ottoman Empire. For centuries, these records were written in a beautiful but tricky script called Ottoman Turkish, which flows from right to left and mixes Turkish words with Arabic and Persian flavors. The problem? The alphabet was swapped out for a modern Latin one in 1928, leaving very few people today who can actually read the old stuff.

To solve this, scientists are building "digital detectives" called Artificial Intelligence. Specifically, they are using a type of AI known as a Vision-Language Model (VLM). Think of a VLM as a super-smart robot that has read millions of books and looked at millions of pictures. It doesn't just "see" a squiggle on a page; it understands that the squiggle is a letter, and that letter is part of a word with a specific meaning. The big question researchers are asking is: Can we teach these robots to read these ancient, messy manuscripts without needing a human expert to sit down and type out every single word for them? If we can, we could unlock history at a speed no human team could ever match.

This paper introduces a new, two-part strategy to teach a robot named "Azra" how to read Ottoman Turkish handwriting. The researchers started with a robot that was already pretty good at reading modern Arabic and English handwriting. They then tried two different ways to teach it the specific quirks of the Ottoman style.

The first method, which they call Azra 1, was like a traditional classroom. The team gathered a small, high-quality collection of 1,306 lines of text that had been carefully checked and typed out by human experts. They also added about 2,000 pages of thesis documents where students had already translated the old script into modern Turkish letters. The robot studied these carefully curated examples and learned to recognize the specific, flowing style of the "Riqa" script (a fast, cursive handwriting used in government offices). When tested on similar handwriting, Azra 1 became incredibly accurate, making fewer than one mistake for every five characters it read. It even beat a popular commercial tool used by historians today.

The second method, Azra 2, was a bold experiment in "self-teaching." The researchers wanted to see if the robot could learn without any human experts typing out the answers. They took thousands of pages of the same old documents and let the first robot (Azra 1) guess what the text said. Then, they used a powerful language tool (an AI called Gemini) to translate those guesses back into the old script. Finally, they used a computer program to check if the robot's guess matched the original page. If the guess was close enough, they kept it as a new training example. This process created a massive dataset of over 23,000 lines of text, all generated by machines with zero human annotation.

The results were fascinating and revealed a trade-off. Azra 1 (the human-taught robot) was the champion when reading the specific style it was trained on. However, when the researchers tested it on a completely different, more formal style of handwriting called "Naskh," it stumbled a bit. It had memorized the specific look of the Riqa script so well that it got confused by anything new.

Azra 2 (the self-taught robot) was a bit less perfect on the familiar script, but it was much more flexible. Because it had seen a wider variety of examples generated through its own bootstrapping process, it handled the new "Naskh" style almost as well as the human-taught robot, and sometimes even better. This suggests that having a huge, diverse pile of data—even if it was made by machines—can make an AI more robust and less likely to get confused by new handwriting styles.

The team also looked closely at the mistakes the robots made. They found that the errors weren't random; they were specific "confusions" caused by the way the machines translated the text back and forth. For example, the AI sometimes mixed up three very similar-looking letters that differ only by the placement of a tiny dot, or it swapped a historical letter for a modern one because the translation tool didn't know the old rules. This is a crucial finding: it means the problem isn't that the robot can't "see" the letters, but that the "translator" in the middle of the process needs to be smarter about the specific history of the language.

In short, the paper shows that we can build powerful tools to read history. While human-verified data is still the gold standard for perfect accuracy on specific styles, a clever, autonomous method can create a massive amount of training data that makes the AI more adaptable to different handwriting styles. The path forward isn't just about feeding the robot more pictures; it's about teaching the robot to understand the subtle, historical rules of the script so it stops mixing up those tricky, dot-less letters.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →