How Robust Are LLMs to Vietnamese Dialects?
This paper introduces VialectBench, the first systematic benchmark evaluating Large Language Models on Vietnamese dialects, revealing that dialectal variations cause measurable performance degradation across multiple tasks and demonstrating that high proficiency in Standard Vietnamese does not guarantee robustness to meaning-preserving regional differences.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are teaching a super-smart robot to understand human language. You show it millions of books, news articles, and websites, all written in a very polished, "textbook" version of a language. The robot gets really good at this, acing every test you give it. But here's the catch: real life isn't a textbook. People don't always speak the "perfect" version. They have accents, they use slang, and they talk differently depending on whether they are from the north, the south, or the middle of the country. This is called a dialect. It's like the same song played on different instruments; the melody (the meaning) is the same, but the sound (the words and grammar) changes.
Scientists have been building these giant language robots, known as Large Language Models (LLMs), to help us write, answer questions, and solve problems. But there's a big question hanging over them: Are they actually smart enough to understand everyone, or are they just really good at understanding the "standard" version of the language? If you ask a robot a question using a local dialect, does it still get the answer right, or does it get confused and start making things up? This matters because if a robot only understands the "perfect" version, it might fail when real people try to use it, leading to frustration or even dangerous misunderstandings.
The Great Vietnamese Dialect Test
A team of researchers decided to put these language robots to the ultimate test. They wanted to see if the models could handle the messy, beautiful reality of Vietnamese dialects. In Vietnam, people speak in many different regional flavors. A word that means "what" in the capital city might be a completely different word in the central provinces, even though everyone understands what is being asked.
The researchers built a special playground called VialectBench. Think of it as a giant obstacle course designed specifically to trip up robots that are too picky about how things are said. They started with 400 standard questions and tasks—like asking a robot to guess someone's mood, solve a logic puzzle, or answer a question based on a story. Then, they hired human experts from six different regions of Vietnam to rewrite those exact same tasks using their local dialects.
The goal was simple: keep the meaning exactly the same, but change the words to sound like a local. For example, instead of asking "What did the doctor tell me to eat?" in the standard way, a person from the central region might ask, "Bác sĩ dặn ăn chi?" using a local word for "what." The robot had to answer both versions. If the robot got the standard version right but failed the dialect version, it meant the robot was brittle—it was easily confused by a change in accent.
The Results: Robots Get Lost in the Middle
When the researchers ran the tests on ten different language models (ranging from small, open-source ones to the massive, powerful GPT-4o), the results were a bit of a wake-up call. No robot was perfect. Every single one of them stumbled when faced with dialects.
On average, the robots' performance dropped by 2.82% when they had to deal with dialects instead of standard Vietnamese. But the drop wasn't the same everywhere. It was like the robots had a specific blind spot.
- The "Central" Problem: The biggest trouble came from the Central dialects (specifically the groups labeled PNT2 and PNT3). When the robots encountered these dialects, their performance took a nosedive. The PNT3 dialect caused the biggest average drop, hurting the robots by 6.17%, while PNT2 dropped them by 4.73%.
- The "Northern" Surprise: Interestingly, the Northern dialect (PNB) was actually the easiest for the robots. In fact, for some models, hearing the Northern dialect actually made them perform slightly better (by 0.42%), likely because the standard Vietnamese the robots were trained on is already very close to the Northern way of speaking.
- The Southern Mix: The Southern dialect (PNN) caused a moderate drop of about 1.81% on average.
Where Did They Fail?
The researchers also looked at what kind of tasks the robots failed at. It turned out that Question Answering (QA) was the most vulnerable. When the robots had to read a story and answer a question in a dialect, they struggled the most, with an average performance drop of over 5%. This suggests that when the question itself is phrased in a local dialect, the robot has a hard time connecting the dots to find the answer in the text.
Another scary finding was the "harmful flip." This happens when a robot gets the answer right in standard Vietnamese but gets it completely wrong in the dialect. The Central dialects caused the most of these flips, with a rate of 6.54% across all models. It's like the robot confidently saying, "I know the answer!" in one voice, and then confidently saying, "The answer is this!" in a different voice, when the two answers are actually opposites.
What This Means
The study showed that just because a robot is a genius at standard Vietnamese doesn't mean it's a genius at all Vietnamese. The researchers found that strong performance on standard language does not guarantee reliable behavior under regional variation.
They also ruled out the idea that these robots are naturally dialect-proof. Even the smartest models, like GPT-4o, which performed the best overall, still showed a small drop in accuracy when dialects were introduced. The study suggests that we can't just assume these tools work for everyone; they need to be tested and improved specifically for the diverse ways people actually speak.
In short, the robots are like tourists who have studied a guidebook perfectly but get lost the moment a local gives them directions in a thick accent. Until we teach them to listen to all the different voices, they might keep missing the mark in the real world.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.