← Latest papers
💬 NLP

VIVID: A Culturally Grounded Benchmark Exposing the Figurative Language Gap in Vietnamese NLP

The paper introduces VIVID, the first culturally grounded benchmark for Vietnamese figurative language, which reveals that current state-of-the-art models, including Vietnamese-specialized ones, significantly struggle with nuanced idioms and proverbs due to literal over-interpretation and a lack of cultural competence.

Original authors: Tu Tran Do, Nhat Ngoc Nguyen, Khanh-Tung Tran, Hoang D. Nguyen, Tu Minh Phuong, Long Hoang Dang

Published 2026-08-05
📖 3 min read☕ Coffee break read

Original authors: Tu Tran Do, Nhat Ngoc Nguyen, Khanh-Tung Tran, Hoang D. Nguyen, Tu Minh Phuong, Long Hoang Dang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot to understand human jokes, sarcasm, and old-fashioned sayings. This isn't just about knowing what words mean; it's about understanding the culture behind them. In the world of computer science, this field is called Natural Language Processing (NLP). Think of NLP as the bridge between human speech and computer code. For a long time, scientists have been building bridges for big, popular languages like English, but many smaller or culturally rich languages have been left on the other side of the river. A key challenge in this field is "figurative language"—phrases where the words don't mean what they literally say. For example, if someone says, "It's raining cats and dogs," a computer that only knows literal definitions might start looking for falling animals! To solve this, researchers need special tests, or "benchmarks," to see if their AI models can actually get the joke or the lesson behind the words, rather than just guessing based on patterns.

Enter VIVID, a new and colorful tool created by researchers to test how well AI understands Vietnamese idioms and proverbs. Think of VIVID as a giant, culturally specific "trivia night" designed specifically for Vietnamese sayings. The researchers gathered 1,636 of these tricky phrases—some are short, rhythmic proverbs full of life lessons, while others are fixed idioms that can't be understood word-for-word. They didn't just throw them into a pile; they organized them like a museum exhibit, tagging each one with "difficulty levels" (like whether it uses ancient words, refers to farming, or relies on sarcasm) and "themes" (like love, criticism, or work).

The team then invited eight of the smartest AI models in the world to take this test. These included both models built specifically for Vietnamese and massive, general-purpose models that speak many languages. They asked the AIs to explain the sayings and to guess what category they belonged to. The results were a bit of a wake-up call. Even the most advanced AI models, including the famous GPT-4o, struggled mightily. On average, they got less than half of the answers right. The study found that the AIs often made silly mistakes, like taking a metaphor too literally (thinking a phrase about a buffalo's skin was actually about a person's personality) or missing the sarcastic tone entirely. Surprisingly, giving the AI a few examples to learn from (a technique called "few-shot prompting") didn't always help; in fact, for the top model, it sometimes made things worse by confusing the AI with the wrong style of answer.

The paper suggests that while these AI models are getting better at general language, they still lack "cultural competence." They are like tourists who can read a map but don't understand the local customs. The researchers conclude that current models are missing the deep, nuanced understanding needed to truly grasp the heart of Vietnamese figurative language. VIVID is now available as a public tool, acting as a mirror to show developers exactly where their AI is failing, so they can build systems that don't just speak the language, but truly understand the culture behind it.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →