← Latest papers
💬 NLP

Figurative and Cultural Knowledge in LLMs: Investigating Cross-Domain Transfer through Fine-Tuning

This paper investigates whether fine-tuning on cultural data enhances figurative language understanding in LLMs, finding that while poetry training improves idiom comprehension in a transferable manner, cultural fine-tuning often degrades proverb interpretation and reveals a complex, non-linear relationship between cultural immersion and figurative language acquisition that cannot be straightforwardly captured through fine-tuning alone.

Original authors: Mena Attia, Mona Diab, Thamar Solorio

Published 2026-08-20
📖 4 min read☕ Coffee break read

Original authors: Mena Attia, Mona Diab, Thamar Solorio

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Language is more than a set of rules for arranging words; it is a living archive of a community's history, values, and shared experiences. This is especially true for figurative language—the idioms, proverbs, and poetry that rely on cultural context to make sense. A phrase like "it's raining cats and dogs" means nothing to someone who has never heard the expression, but to a native speaker, it instantly conjures a specific image of heavy rain. For artificial intelligence, which learns by processing vast amounts of text, mastering this kind of language is a significant hurdle. While modern AI models can read and write fluently, they often struggle to understand the deeper, non-literal meanings that are woven into the fabric of a culture. The question researchers have long asked is whether teaching an AI about a culture's daily life and traditions would help it understand its metaphors, or if learning to interpret metaphors would help it grasp cultural facts.

A team of researchers set out to test this connection using Arabic, a language rich with diverse dialects and a deep literary tradition. They worked with four different large language models, which are the powerful computer systems that drive many modern AI applications. To see if cultural knowledge and figurative understanding are linked, the researchers conducted a series of controlled experiments. They took these models and gave them extra training, or "fine-tuning," on specific types of data. In one group, they fed the models thousands of examples of cultural knowledge, such as facts about social customs, regional practices, and everyday life across the Arab world. In another group, they trained the models on figurative language, using datasets filled with proverbs, idioms, and poetry. They also included a control group trained on general Arabic grammar to ensure that any changes were due to the specific content of the training and not just the act of learning more Arabic.

The results revealed a complex and somewhat surprising picture. The researchers found that teaching a model about culture did not automatically make it better at understanding idioms or proverbs. In fact, for models that were already specialized in Arabic, training them on cultural data sometimes made their performance worse. These models, which had likely already absorbed a great deal of cultural knowledge during their initial training, seemed to become confused when exposed to a narrower, curated set of cultural facts. Conversely, training models on cultural data did not help them answer general questions about cultural facts either. The two types of knowledge—cultural facts and figurative meaning—did not seem to support each other in the way the researchers had hoped.

However, there was one clear exception that stood out. When the researchers trained the models specifically on poetry, the models became significantly better at understanding idioms. This improvement was measurable and statistically reliable, with a specific model showing a notable jump in accuracy. Crucially, this gain did not happen when the models were trained on general Arabic text or grammar, which suggests that the improvement came from the poetic content itself, not just from seeing more Arabic words. Poetry, with its reliance on imagery, rhythm, and non-literal meaning, appears to have taught the models a broader skill: the ability to recognize and interpret meaning that goes beyond the literal definition of words. This skill transferred successfully to idioms, even though poetry and idioms are different forms of expression.

The study also highlighted that the results depended heavily on which model was being used. The models that were designed specifically for the Arabic language started with a higher baseline of knowledge and often regressed, or performed worse, after being fine-tuned on new data. This suggests they had already learned as much as they could from the available training materials. In contrast, multilingual models, which are trained on many languages, showed more room to learn and improved more consistently when exposed to new data. When the researchers looked closely at the errors, they found that the training tended to reinforce knowledge about everyday, experiential things like food, games, and celebrations, while sometimes weakening the models' grasp of historical facts or specific political details.

Ultimately, the research suggests that the relationship between culture and figurative language is not a simple switch that can be flipped by feeding an AI more data. While the two are deeply connected in human experience, teaching an artificial intelligence one does not guarantee it will master the other. The only reliable bridge found in this study was from poetry to idioms, indicating that exposure to the artistic and non-literal side of a language can help a model navigate its metaphorical landscape. For now, the path to truly culturally aware artificial intelligence remains uneven, requiring more than just a simple transfer of knowledge from one domain to another.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →