Polite on the Surface, Wrong in Practice: A Curated Dataset for Fixing Honorific Failures in Multilingual Bangla Generation
This paper introduces BLADE, a curated dataset of 4,196 interaction pairs designed to address honorific and pragmatic failures in multilingual Bangla generation, demonstrating that fine-tuning open-weight models on this resource significantly improves structural fidelity and cultural alignment.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Core Problem: The "Polite but Clueless" Robot
Imagine you ask a very smart, well-read robot to write a formal letter to your school principal asking for a few days off. The robot writes something that looks perfect at first glance. It has big words, it's grammatically correct, and it flows well.
But then, you read closer. The robot starts the letter saying, "Dear Sir," but halfway through, it switches to talking to the principal like they are your best friend, using slang and informal pronouns. It might say, "Hey, can you give me a break?" right after saying, "I respectfully request your permission."
To a native speaker, this is a disaster. It's like wearing a tuxedo to a funeral but then doing a backflip on the way in. The robot sounds fluent, but it fails the most important test: social competence.
The paper argues that current AI models (LLMs) are great at learning vocabulary but terrible at learning cultural rules, especially for languages like Bangla where the way you speak changes based on who you are talking to (honorifics).
The Solution: The "BLADE" Dataset
To fix this, the researchers built a special training manual called BLADE (BangLa Applications and DialoguEs).
Think of existing AI training data as a giant library of random books, news articles, and internet comments. It's huge, but it's messy. If you want to learn how to write a formal letter, you might find a few examples, but they might be mixed with casual text or contain errors.
BLADE is different. It's like a strict, expert-led writing workshop.
- The Content: It contains 4,196 carefully crafted examples of formal applications (like leave requests) and dialogues.
- The Teachers: These weren't just written by computers. They were curated by humans, checked against government textbooks, and verified by language experts to ensure the "politeness levels" (honorifics) were perfect.
- The Rule: Every single example teaches the AI that in a formal setting, you must always use the "respectful" version of "you," never the "casual" version, and you must follow a specific structure (Date → Recipient → Subject → Body → Sign-off).
The Experiment: Teaching the Robot the Rules
The researchers took several popular AI models (some huge, some small) and gave them a choice:
- The "Zero-Shot" Test: Ask the AI to write a letter without any special training. (Result: The AI sounded fluent but made rude mistakes, like mixing formal and informal language).
- The "Fine-Tuning" Test: Feed the AI the BLADE dataset to learn the specific rules of formal Bangla writing.
The Results:
- Before Training: Even the biggest, most expensive AI models failed to write a usable letter. They got low scores because they didn't understand the social rules.
- After Training: The models improved dramatically.
- The "Small" Winner: A tiny model (only 1.5 billion parameters) trained on BLADE actually performed better than the massive, untrained models.
- The Analogy: It's like taking a small, local chef who knows the exact recipe for a traditional dish and giving them a perfect cookbook. They will cook a better meal than a world-famous chef who has never seen that specific recipe before.
Key Takeaways
- Data Quality > Model Size: For languages with complex social rules (like Bangla), having a massive brain isn't enough. You need the right training data. A small model with high-quality, culturally specific data beats a giant model with generic data.
- The "Pragmatic Gap": There is a huge difference between a model that can speak a language and a model that can use it correctly in real life. The paper shows that without specific training, AI fails at the "real life" part.
- Structure Matters: The AI didn't just learn words; it learned the shape of a formal document. It learned that a formal letter needs a date at the top and a specific sign-off at the bottom, not just random paragraphs.
What the Paper Does Not Claim
- It does not claim this fixes all AI problems.
- It does not claim this works for informal slang or casual texting (the dataset was focused on formal, institutional writing).
- It does not claim the models are now perfect; they are just much better at this specific task than before.
In a nutshell: The paper says, "If you want an AI to write a formal letter in Bangla, don't just give it a bigger brain. Give it a better teacher and a strict rulebook."
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.