How do datasets, developers, and models affect biases in a low-resourced language?: The Case of the Bengali Language
This paper empirically demonstrates that sentiment analysis models for the low-resource Bengali language exhibit persistent gender, religious, and nationality-based biases regardless of whether they are built on multilingual or monolingual architectures, highlighting how the combination of diverse datasets and pre-trained models can introduce inconsistencies that exacerbate epistemic injustice in sociotechnical systems.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a robot to understand human feelings (like happiness, sadness, or anger) by showing it thousands of sentences written in Bengali. You want the robot to be fair to everyone, no matter if they are a man or a woman, Hindu or Muslim, or from Bangladesh or India.
This paper is like a "report card" for 38 different versions of that robot. The researchers asked: Where does the robot get its unfairness? Is it from the books it read (the datasets), the teacher who taught it (the developers), or the brain it was built with (the pre-trained model)?
Here is the breakdown of their findings using simple analogies:
1. The Setup: The Robot, The Library, and The Teachers
- The Robots (Models): The researchers started with two "base brains." One was a generalist (mBERT) that knows 104 languages but isn't an expert in any of them. The other was a specialist (BanglaBERT) that only knows Bengali and was trained specifically on Bengali books.
- The Libraries (Datasets): They found 19 different collections of Bengali sentences (datasets) from the internet. Think of these as different libraries. Some libraries were filled with news, some with social media posts, and some with novels.
- The Teachers (Developers): These libraries were organized by different people. The researchers tried to find out who these teachers were (their gender, religion, and nationality) to see if the teachers' personal biases were being passed down to the robots.
2. The Experiment: The "Taste Test"
The researchers gave the robots a special "taste test." They fed them pairs of sentences that meant the exact same thing but used different words to signal identity.
- Example: One sentence used a word for "water" common in India, and the other used a word for "water" common in Bangladesh.
- The Goal: If the robot is fair, it should give both sentences the same "feeling score." If it gives one a high score (positive) and the other a low score (negative), it is biased.
3. The Results: What Did They Find?
A. The Robots Are Biased (The "Echo Chamber" Effect)
The study found that most of the robots were unfair.
- Gender: 61% of the robots liked sentences about men more than women. Only 24% liked women more.
- Religion: 61% of the robots favored Muslim identities, while 24% favored Hindu identities.
- Nationality: 50% of the robots liked Bangladeshi identities more, while 26% preferred Indian identities.
B. The Teachers Didn't Matter (The "Ghost in the Machine")
You might think, "If the teacher is a man, the robot will be sexist." The researchers checked this, but they found no clear link.
- Even though most of the dataset creators were men, Muslims, and from Bangladesh, the robots didn't just copy their specific biases.
- The Takeaway: The bias didn't come from the specific person who typed the data. It came from the interaction between the data and the robot's brain.
C. The Brain Matters More Than the Library (The "Specialist vs. Generalist")
This was the most surprising finding.
- The Specialist Robot (BanglaBERT) was generally fairer and less biased than the Generalist Robot (mBERT).
- The Analogy: Imagine a generalist robot is like a tourist who tries to understand a culture by reading a few travel brochures in many languages. They get the basics but miss the nuances. The specialist robot is like a local who grew up speaking the language; they understand the subtle cultural differences better.
- However, no library was perfect. Even the best datasets had biases in some areas. If a dataset was fair regarding gender, it might be unfair regarding religion.
4. The Big Picture: It's a Chain Reaction
The paper argues that bias isn't just a "bug" in one piece of code. It's a chain reaction.
- Think of it like a relay race.
- The first runner (the pre-trained model) starts with some inherent biases.
- The second runner (the dataset) adds their own flavor.
- When they pass the baton (fine-tuning), the final result is a mix of both.
- Sometimes, a "good" dataset can fix a "bad" model. Sometimes, a "bad" dataset can ruin a "good" model.
5. Why This Matters (The "Epistemic Injustice")
The authors use a fancy term called Epistemic Injustice. In simple terms, this means the robot is unfairly discrediting certain people's voices.
- If a robot thinks a sentence written by a Bangladeshi person is "negative" just because they used a specific local word for "water," the robot is saying, "Your way of speaking is wrong or bad."
- This is unfair because it treats one group's language as the "standard" and others as "incorrect," reinforcing old social divides.
Summary
The paper concludes that to fix bias in low-resource languages like Bengali, we can't just blame the people who made the data. We have to look at the whole system:
- Don't just use one-size-fits-all models: Specialized models (like BanglaBERT) seem to handle local culture better than general ones.
- No dataset is perfect: Every library has some bias, so we need to be careful about which one we pick.
- The combination is key: The bias comes from how the model and the data mix together, not just from one or the other.
The researchers are calling for more "audits" (checks and balances) to make sure these digital tools don't accidentally hurt the communities they are supposed to serve.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.