← Latest papers
💬 NLP

5-Dialects-BN: Unmasking the Impact of Transliteration on Bangla Dialectal LLMs

This paper introduces 5-Dialects-BN, the first multi-annotation benchmark aligning Romanized transliteration with dialectal, standard, and English texts alongside subjectivity labels for five major Bangla regional varieties, aiming to bridge the resource gap and enable principled evaluation of dialect-aware Large Language Models.

Original authors: Md Mahir Jawad, Galib Mahmud Jim, Rafid Ahmed, Mir Sazzat Hossain, Md Fahim, Md Farhad Alam Bhuiyan

Published 2026-09-10
📖 4 min read☕ Coffee break read

Original authors: Md Mahir Jawad, Galib Mahmud Jim, Rafid Ahmed, Mir Sazzat Hossain, Md Fahim, Md Farhad Alam Bhuiyan

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Language is not a single, solid block; it is a living landscape that shifts and changes as people move across regions. In the world of artificial intelligence, large language models have become incredibly skilled at understanding and generating text, but they have largely learned from a very specific, formal version of languages. When these digital minds encounter the messy, vibrant reality of how people actually speak in their local neighborhoods, they often stumble. This is particularly true for Bengali, a language spoken by hundreds of millions of people across Bangladesh and India. While the standard, formal version of the language is well-documented, the many regional dialects spoken in cities like Chittagong, Sylhet, and Barisal have been largely invisible to the technology. People in these regions often type in a mix of their local dialect and English letters, a practice known as transliteration, creating a unique digital voice that existing systems struggle to hear.

A team of researchers set out to map this uncharted territory by creating a new resource called 5-DIALECTS-BN. They gathered 6,000 real sentences from social media and online forums, capturing the natural way people speak in five distinct regional dialects: Chittagong, Barisal, Noakhali, Sylhet, and Rangpur. For every sentence, they did not just record the words; they built a complete bridge between the different ways people express the same thought. Each entry includes the original text in the local dialect, a version written in English letters (transliteration), a translation into standard Bengali, a translation into English, and a label indicating whether the sentence expresses a personal opinion or a simple fact. This careful, human-verified work ensures that the data reflects the true nuances of the language, from unique slang words to subtle shifts in pronunciation that change the meaning of a sentence.

The researchers then put this new dataset to the test, asking seven different large language models to perform three specific tasks: translating the dialect into English, deciding if a sentence is an opinion or a fact, and converting the local dialect back into standard Bengali. They tested the models in several ways, including giving them no examples, showing them a few examples to learn from, and using a method where the models were allowed to "think" through the problem step-by-step before answering. They also tried a technique called fine-tuning, where they taught the models using just a small number of examples—160 sentences for each dialect—to see if a little bit of direct instruction could help them understand better.

The results revealed a surprising and critical flaw in how these powerful models handle language. The researchers found that when the input text was written in English letters (transliterated), the models performed significantly worse across the board, regardless of how smart the model was or how many examples they were given. This happened because the process of converting local sounds into English letters often loses crucial information. Different sounds in the local dialect can look identical when written in English letters, creating confusion that the models cannot resolve. It is as if a listener is trying to understand a conversation through a wall that muffles specific frequencies; the words are there, but the distinct qualities that make them clear are gone. Even when the models were given a small amount of direct training to help them, the gap between understanding the native script and the English-letter version remained wide, proving that the loss of information in the translation to English letters is difficult to recover.

However, the study also offered a clear path forward. When the researchers used the fine-tuning method with just 160 examples per dialect, the open-source models improved dramatically. They began to outperform even the most advanced, closed-source models that had never seen this specific data before. This suggests that the biggest barrier to understanding these dialects is not a lack of computing power or intelligence, but simply a lack of specific, aligned data. The models are capable of learning these dialects if they are given the right examples. The study concludes that while current technology struggles with the ambiguity of transliterated text, providing it with high-quality, human-verified examples allows it to bridge the gap. This work lays a foundation for building systems that can truly understand the diverse, rich ways people communicate in their daily lives, moving beyond the formal standard to embrace the full spectrum of the language.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →