← Latest papers
💬 NLP

BANGLASOCIALBENCH: A Benchmark for Evaluating Sociopragmatic and Cultural Alignment of LLMs in Bangladeshi Social Interaction

This paper introduces BANGLASOCIALBENCH, the first benchmark comprising 1,719 culturally grounded instances to evaluate sociopragmatic competence in Bangla, revealing that current large language models systematically fail to navigate the language's complex social hierarchies, kinship reasoning, and interactional norms despite their multilingual fluency.

Original authors: Tanvir Ahmed Sijan, S. M Golam Rifat, Pankaj Chowdhury Partha, Md. Tanjeed Islam, Md. Musfique Anwar

Published 2026-03-18
📖 4 min read☕ Coffee break read

Original authors: Tanvir Ahmed Sijan, S. M Golam Rifat, Pankaj Chowdhury Partha, Md. Tanjeed Islam, Md. Musfique Anwar

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: Fluency vs. "Social Smarts"

Imagine you hire a robot butler who speaks perfect English. It has a massive vocabulary and can recite poetry. However, if you ask it to tell a joke at a funeral, or if it calls your strict boss by their first name, the robot is a disaster. It knows the words, but it doesn't understand the rules of the room.

This paper argues that Large Language Models (LLMs) like the ones powering chatbots are currently like that robot butler. They are fluent in many languages, including Bangla (spoken in Bangladesh), but they often fail at sociopragmatics—the art of knowing how to speak to whom, in what situation, to sound polite and culturally correct.

The Problem: The "Three-Word" Trap

In English, we mostly have one word for "you." In Bangla, there are three different words for "you," and choosing the wrong one is like wearing a tuxedo to a beach party or a swimsuit to a wedding.

  1. Apni: The "Respectful You." Used for elders, bosses, or strangers. (Like saying "Sir/Ma'am").
  2. Tumi: The "Friendly You." Used for peers, friends, or slightly younger people. (Like saying "Hey buddy").
  3. Tui: The "Intimate/Childish You." Used for very close friends, siblings, or children. (Like saying "Little one" or "Buddy" in a very casual way).

The Mistake: If an AI talks to an elderly stranger using Tui, it sounds incredibly rude. If it talks to a 5-year-old using Apni, it sounds weirdly stiff and distant. The paper found that current AI models are terrible at guessing which "You" to use based on the context.

The Solution: BANGLASOCIALBENCH

The researchers built a test drive (a benchmark) specifically for Bangla social skills. Think of it as a driving test for cultural etiquette.

Instead of asking the AI, "What is the capital of Bangladesh?" (a fact), they ask it:

"You are an elderly man at a metro station. A teenage boy is cutting in line. How do you politely ask him to wait his turn?"

The AI has to choose the right word and tone.

  • Wrong Answer: "Hey kid, get in line!" (Too rude).
  • Right Answer: "Son, please stand in the line." (Polite, using the right kinship term).

The test covers three main areas:

  1. Address Terms: Choosing the right "You" (Apni/Tumi/Tui) or title (Uncle, Sister, Sir).
  2. Kinship Reasoning: Figuring out family trees. In Bangla, "Uncle" isn't just "Uncle." There is a specific word for your father's younger brother vs. your mother's brother. The AI often gets confused here, mixing up Hindu and Muslim family terms.
  3. Social Customs: Knowing how to behave. For example, if a guest says "I'm full," a Bangladeshi host will insist they eat more (hospitality). An AI might just say, "Okay, no problem," which misses the cultural script.

What They Found: The "Over-Polite" Robot

The researchers tested 12 different AI models (including big names like GPT-4, Gemini, and Llama). Here is what they discovered:

  • The "Safe" Strategy: The models are terrified of being rude. So, they almost always choose the most formal, stiff option (Apni). They would rather sound like a robot at a funeral than a friend at a party.
  • The "Downward" Failure: The AI is okay when a younger person talks to an elder (because that's always formal). But when an elder talks to a younger person, the AI gets confused. It doesn't know when to relax and use a friendly tone.
  • The Religious Mix-Up: The models often confuse Hindu and Muslim family terms. If the story implies a Hindu family, the AI might accidentally use a Muslim term, showing it doesn't truly understand the cultural nuance, just the words.
  • The "One-Size-Fits-All" Brain: Even when humans agree that two different words could be correct in a situation, the AI picks one and sticks to it with 100% confidence. It lacks the human ability to say, "Well, you could say it this way or that way."

The Takeaway

This paper is a wake-up call. We can't just teach AI more facts about a culture. We have to teach it social intuition.

Currently, AI is like a tourist who has memorized a phrasebook but doesn't understand the local customs. It can say "Hello," but it doesn't know how to say hello to a grandmother versus a business partner. Until we fix this, AI will always feel a little "off" when interacting with real people in high-context cultures like Bangladesh.

In short: The AI knows the dictionary, but it hasn't learned the dance.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →