Wisdom in Unity: The Role of Multilingual Training in Figurative Language Identification in Proverbs
This paper demonstrates that multilingual supervision, particularly when incorporating diverse figurative forms like Culture Specific and Moral/Advisory proverbs, significantly enhances figurative language identification performance and suggests a shift toward multidimensional, concept-level frameworks beyond metaphor-centric approaches.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a robot to understand human jokes, riddles, and old sayings. This isn't just about translating words from one language to another; it's about understanding the hidden meaning behind them. In the world of computer science, this is called "figurative language." Think of it like a secret code where "it's raining cats and dogs" doesn't mean animals are falling from the sky, but that the rain is heavy. Proverbs are the ultimate test for these robots because they are short, packed with cultural wisdom, and often make no sense if you take them literally. For a long time, scientists tried to teach computers to spot these sayings using only one language at a time, like studying English jokes in an English-only classroom. But the real world is messy and multilingual, so researchers wanted to know: if we show a robot sayings in many different languages at once, does it get smarter? Does seeing a proverb in French, Japanese, and Arabic help it understand the proverb in English better?
This paper, titled "Wisdom in Unity," dives into that exact question using a massive collection of 742 proverb concepts translated into seven different languages (Arabic, English, French, German, Russian, Japanese, and Spanish), totaling over 6,700 examples. The researchers didn't just ask if a proverb is "figurative" or "literal"; they broke it down into four specific flavors of meaning: Metaphorical (using images or symbols), Moral/Advisory (giving a lesson or warning), Cause–Effect (linking an action to a result), and Culture-Specific (needing deep local knowledge to get the joke). They tested five different computer models, ranging from standard encoders to advanced "instruction-tuned" large language models (LLMs), to see how well they performed when trained with different amounts of these translated examples.
Here is what they discovered, and it's a bit like finding the perfect recipe for a multilingual smoothie. First, they found that you don't need to feed the robot every single translated proverb to get great results. In fact, feeding the models about 50% of the translated training data was enough to reach near-perfect performance. Adding more data beyond that point (going from 50% to 100%) didn't help much, suggesting that the models hit a "sweet spot" where they had learned enough to generalize.
Second, and perhaps most surprisingly, the type of proverb matters a lot. The researchers found that the "Culture-Specific" proverbs—the ones that are the hardest to understand because they rely on local history or religion—actually saw the biggest improvement when the models were trained with multilingual data. It's as if seeing the proverb in different cultural contexts helped the robot finally crack the code of what makes it unique. Meanwhile, the more common types, like metaphors, didn't get as big of a boost from the extra languages.
Finally, the study showed that no single "flavor" of proverb is enough on its own. The models performed best when they were trained on a mix of all four types. It turns out that while some models might be great at spotting moral lessons and others at spotting cause-and-effect, combining all these different perspectives creates the strongest overall understanding. The authors suggest that instead of just looking for metaphors, future AI needs to be trained to recognize these different, complementary layers of meaning to truly understand the wisdom hidden in our global collection of proverbs.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.