Preserving Myanmar’s Ancient Heritage: A Novel Dataset and ResNet Architecture for Pyu Character Recognition
This paper addresses the critical lack of computational resources for Myanmar's ancient Pyu script by introducing the first publicly available, systematically curated handwritten dataset of 33 consonants and validating a modified ResNet-18 architecture as a foundational benchmark for future digital preservation and epigraphic research.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a world where ancient history is trapped behind a glass wall that computers simply cannot see through. For decades, scientists have used powerful tools called "computer vision" to teach machines how to recognize images—like spotting a cat in a photo or reading a street sign. They also use "deep learning," a method where computers learn by looking at thousands of examples, much like a student studying flashcards until they get every answer right. But there is a huge gap in this digital library: while we have great tools for modern languages, the ancient scripts of the world often remain invisible to these machines. This is a problem because many of these old writings are fading away, and if we can't teach computers to read them, we might lose the stories, laws, and prayers of our ancestors forever. The question isn't just about technology; it's about saving a part of human memory that is currently stuck in a format computers don't understand.
This paper tackles that exact problem for the ancient Pyu script, a writing system used in Myanmar over a thousand years ago. The researchers faced a "chicken and egg" situation: to teach a computer to read Pyu, you need a massive library of examples (a dataset), but because the script is so old and rare, no such library existed. To solve this, the author, Khant Sint Heinn, didn't just write code; they became a digital scribe. They physically hand-wrote every single one of the 33 Pyu consonants, creating 25 different versions of each letter to capture the natural wobbles and styles of human handwriting. This resulted in 825 original images.
But 825 pictures isn't enough to train a super-smart computer. So, the author used a clever trick called "data augmentation." Think of this like a photocopier that doesn't just copy a page, but also rotates it, blurs it slightly, stretches it, and tilts it to look like it was written on a bumpy, weathered stone tablet. By applying these five different "distortions" to every single image, they turned their 825 original drawings into a massive library of 4,950 images. They then fed this new library into a modified version of a famous computer brain architecture called ResNet-18.
The results were incredibly promising. In a controlled test where the computer had to identify these clean, high-contrast drawings, the model learned almost instantly. After just three rounds of studying, the computer achieved a perfect score, correctly identifying 100% of the unseen test images. The confusion matrix—a chart that shows if the computer mixed up one letter for another—was a perfect diagonal line, meaning it never made a mistake on these specific samples.
However, the paper is very careful not to call this a "finished" solution for reading ancient history. The authors explain that while the computer aced the test on their clean, hand-drawn images, real-world ancient inscriptions are much messier. Real stone tablets found in archaeological sites are cracked, covered in dirt, eroded by centuries of rain, and often photographed in poor lighting. The paper argues that while this study proves the computer can learn the shape of the letters, it hasn't yet proven it can read them off a crumbling stone wall. This work is described as a foundational benchmark—a solid first step that proves the data exists and the letters are distinct enough to be learned. The authors suggest that future work will need to use "transfer learning" to take this knowledge and apply it to the noisy, damaged reality of actual historical artifacts. For now, they have successfully built the digital dictionary that was missing, opening the door for future researchers to finally teach computers to read Myanmar's ancient past.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.