AI-assisted Egyptian hieroglyphic reading in scholarly practice: workload, usability, and acceptance in a pilot feasibility study
This pilot feasibility study evaluates the impact of the AI tool PyThoth on Egyptologists' workload, usability, and acceptance during hieroglyphic reading, revealing that while the tool improves task completion rates, it does not necessarily enhance accuracy and elicits complex, experience-dependent shifts in user perception.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Ancient Egyptian writing is a puzzle that has challenged scholars for centuries. Unlike modern alphabets where a single letter usually stands for one sound, these hieroglyphs are a complex mix of pictures that can represent sounds, ideas, or even whole words. A single sentence might be read in different directions, and the same word could be spelled in several different ways depending on the time period or the scribe. For a human expert, deciphering these texts requires years of training to recognize thousands of unique symbols, understand their grammatical rules, and piece together meanings from damaged stone surfaces. It is slow, meticulous work that forms the backbone of Egyptology, the study of ancient Egypt.
Recently, computers have become incredibly good at recognizing patterns, leading to the hope that artificial intelligence could help speed up this process. Imagine a tool that looks at a photo of a stone carving, identifies the symbols, suggests how to spell them out in our alphabet, and even offers a translation. This is the promise of AI in the humanities. However, while these tools are being built and tested in laboratories, no one had really asked how actual experts and students feel about using them in their daily work. Does the technology make the job easier, or does it create new problems? Does it help students learn, or does it stop them from learning the necessary skills? To find out, a team of researchers set up a small, real-world test to see how people actually interact with this technology.
The researchers invited nine people to participate, ranging from undergraduate students just starting their studies to seasoned professional Egyptologists. They gave everyone a piece of a real ancient Egyptian stone monument, known as a stela, which had been broken into four smaller sections. The participants were asked to translate these sections into English. Half of the time, they worked alone, using only their knowledge and standard reference books. The other half of the time, they used a new computer program called PyThoth. This program acted as a digital assistant: it would look at the image of the hieroglyphs, suggest what each symbol was, propose a spelling, and offer a translation. Crucially, the computer did not do the work for them. The human had to check every suggestion, correct any mistakes, and decide whether to accept the final result. The researchers wanted to see if this partnership changed how much mental effort the people felt they were using, how easy the tool was to handle, and whether they would want to use it again.
The results were not a simple story of technology saving time or making everything perfect. Instead, the experience depended entirely on the person using the tool and the specific text they were reading. For some participants, the AI assistant made the work feel lighter and more manageable. For others, it felt like a burden that slowed them down or confused them. The study found that the tool did not necessarily make the final translation more accurate than a human working alone. In fact, when the humans finished their work, the accuracy was roughly the same whether they had help or not. The main difference was that the AI seemed to encourage people to finish the task. When working without help, some participants gave up partway through, leaving their translations incomplete. With the AI, they were more likely to push through to the end, even if the final result was not significantly better.
A key discovery was that people's feelings about the tool changed in surprising ways based on their expectations. Some students entered the experiment with high hopes, thinking the AI would work like the smart chatbots they use every day. When they found that the tool required them to click through many steps to fix a single mistake, they became frustrated and felt the tool was harder to use than they expected. Their opinion of the technology dropped sharply after they tried it. On the other hand, a few participants who had low expectations or no experience with such tools were pleasantly surprised. They found that the assistant helped them get through the tedious parts of the work, and their opinion of the tool went up. This suggests that the value of the technology is not just in how smart the computer is, but in how well its design matches what the user expects and needs.
The researchers also noticed that the tool influenced how people wrote their answers. When using the AI, different participants started to write their translations in a very similar way, often copying small errors or specific formatting choices that came from the computer. This happened even though they were working independently. It showed that the tool was shaping their output, sometimes in ways they didn't notice. One expert who tried the system noted that the computer could produce a correct translation even if it had misidentified the symbols in the middle of the process. This raised a subtle concern: a fluent, confident-looking result might hide a flawed understanding of the text, potentially misleading a reader who trusts the computer too much.
Perhaps the most important finding was not about the technology itself, but about how to study it. The researchers expected everyone to sit down and complete the whole experiment in one go. Instead, the reality was much messier. Some people finished one part and stopped. Others came back days or weeks later to finish the rest. Some only filled out the questionnaires without doing the translation work. The study concluded that in the world of specialized humanities research, this kind of scattered, uneven participation is normal. Future studies should not treat these incomplete attempts as failures or lost data, but rather as a natural part of how experts engage with new tools. The path to understanding how AI fits into ancient history is not a straight line, but a complex journey where the human mind, the ancient text, and the machine all interact in unique ways. The technology is not a magic wand that solves everything, but a partner that changes the nature of the work, for better or worse, depending on how it is used.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.