Evaluative Cultural Hallucination in Large Language Model Assessment of Tenor in Scenic Area Public Sign Translations
This study reveals that large language models frequently exhibit "evaluative cultural hallucination" by significantly overestimating the pragmatic and socio-cultural appropriateness (Tenor) of scenic area public sign translations, often mistaking surface formality or recognizable references for genuine cultural suitability.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
When you walk through a historic temple or a sacred burial ground in China, the signs guiding you are more than just directions. They are a bridge between cultures, translating not only words but also the deep respect, history, and solemnity required to speak about a place where ancestors rest or deities are honored. This is the realm of public-sign translation, a task that demands a delicate balance between clarity and cultural sensitivity. For decades, experts have judged these translations by looking at "tenor," a concept that describes the relationship between the writer and the reader. It asks whether the tone is right: is it too casual for a sacred site? Is it too stiff for a nature trail? Does it show the proper level of respect for a historical figure? As cities and tourist sites grow, the need to check thousands of these signs has outpaced the ability of human experts to review them all, leading many to wonder if artificial intelligence can step in to help.
Large language models, the powerful computer systems behind modern chatbots, have shown they can read and write with impressive fluency. They can spot grammar mistakes and suggest smoother phrasing. But a new study from researchers at the University of Malaya asks a harder question: can these machines understand the invisible weight of culture? Specifically, they wanted to know if an artificial intelligence would mistakenly praise a translation that sounds polite and formal but actually misses the deep cultural meaning required for a sacred or historical site. The researchers gathered 116 bilingual signs from eight major tourist attractions in Xuzhou, China, ranging from nature guides to descriptions of ancient tombs and Buddhist temples. They asked both human experts and a sophisticated AI model to grade the English translations based on how well they captured the right tone and respect.
The results revealed a significant gap between human intuition and machine judgment. While the human experts gave many of the translations low scores for failing to capture the necessary cultural respect, the AI model consistently gave them high marks. In nearly half of the cases examined, the AI rated a translation as high quality even when experts said it was culturally inadequate. The problem was most severe with signs related to religion and rituals. For these signs, the AI was wrong about the quality of the translation in 65 percent of the cases. The researchers found that the AI was not simply making random errors; it was following a specific, predictable pattern of misunderstanding. It mistook the surface appearance of a text for its deep cultural value.
When the researchers looked closely at why the AI gave these high scores, they found it was relying on four specific shortcuts. First, the model often praised a translation simply because it sounded formal or institutional, like something you might see in a museum. It assumed that if the words sounded official, the tone was correct, even if the specific cultural respect was missing. Second, the AI was fooled by recognizable names. If the translation kept the name of a historical figure or a religious object, the model assumed the job was done, ignoring whether the surrounding words treated that figure with the proper reverence. Third, the model sometimes confused general politeness with specific cultural solemnity. It would see words that sounded respectful in a broad sense and decide the translation was perfect, failing to notice that the specific ritual nuances were lost. Finally, the AI was too tolerant of missing information. It would give a high score to a translation that was short and easy to read, even if it had cut out important details about rituals or history that were essential to the original meaning.
The study suggests that this is not just a case of the machine being slightly off, but a form of what the authors call "evaluative cultural hallucination." This does not mean the machine is making up facts; rather, it is hallucinating that a translation is culturally appropriate when it is not. It sees the surface features—formal words, recognizable names, or a polite tone—and convinces itself that the deep cultural requirements have been met. This is particularly dangerous for signs in religious or ritual settings, where getting the tone wrong can be seen as disrespectful or even offensive. The researchers found that the AI was much better at judging signs about nature or general information, where the tone is less complex. But when the text involved sacred spaces, burial sites, or ancient rituals, the machine's confidence became a liability.
This research does not say that artificial intelligence cannot help with translation. Instead, it warns that we cannot simply trust the machine's score, especially when culture is at stake. The AI is excellent at checking if a sentence is grammatically correct or if a name is spelled right, but it struggles to understand the invisible rules of respect and hierarchy that govern how we speak about the sacred. The study concludes that for signs in tourist areas, particularly those dealing with history and religion, human experts must remain in the loop. The machine can offer a first look, but its reasoning must be checked by a person who understands that a formal tone is not the same as a respectful one, and that a recognizable name does not guarantee a culturally accurate translation. Until machines can truly understand the weight of culture, the most reliable way to ensure these signs honor the places they describe is to keep a human hand on the wheel.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.