← Latest papers
💬 NLP

AmchiBias: Measuring Stereotypical Bias in Goan Identity Groups with a Minimal Pair Dataset in English and Konkani

This paper introduces AmchiBias, the first benchmark for measuring socio-cultural stereotypical bias in Goan identity groups using minimal pairs in English and Konkani, revealing that current multilingual models lack genuine Goan cultural competence and often reflect pan-Indian biases rather than hyperlocal realities.

Original authors: Michelle Barbosa, Sebastian Padó, Franziska Weeber

Published 2026-06-16
📖 5 min read🧠 Deep dive

Original authors: Michelle Barbosa, Sebastian Padó, Franziska Weeber

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a group of very smart, well-read robots (AI models) that have read billions of books and websites. You want to know if these robots have picked up on the same silly, unfair stereotypes that humans sometimes hold about different groups of people.

Most researchers have only tested these robots on big, national-level stereotypes (like "All people from Country X are lazy"). But the authors of this paper, Michelle Barbosa and her team, asked: What about the tiny, specific stereotypes that exist in just one small corner of the world?

They chose Goa, a small, colorful state in India with a unique mix of history, languages, and cultures. They built a special test called AmchiBias (which means "Our Bias" in the local language, Konkani) to see if these robots understand Goan life or if they are just guessing.

Here is how they did it and what they found, explained simply:

1. The Test: A "Spot the Difference" Game

The researchers created a game of "Minimal Pairs." Imagine two sentences that are identical twins, except for one word:

  • Sentence A: "The Fishermen work tirelessly to support their families."
  • Sentence B: "The Landlords work tirelessly to support their families."

In Goan culture, one of these groups might be stereotyped as "hardworking" while the other is stereotyped as "lazy" (or vice versa). The researchers asked the AI: "Which sentence sounds more natural or true to you?"

They created 313 of these pairs covering eight different topics:

  • Caste: Different social groups (like Bamon or Chardo).
  • Language: People who speak English vs. Konkani vs. Marathi.
  • Jobs: Fishermen, bakers, landlords, politicians.
  • Religion: Catholics, Hindus, Muslims.
  • Origin: Locals vs. Migrants vs. Tourists.
  • Region: People from the North vs. South of Goa.
  • Age: Youth vs. Elderly.
  • Gender: Men vs. Women.

They tested the robots in two languages: English (the language of privilege and education in Goa) and Konkani (the native language of the people).

2. The Results: The "Bilingual" Robot Problem

The researchers tested five different AI models. Here is what happened:

In English: The Robots "Know" the Stereotypes
When the robots read the test in English, they acted like they were steeped in Indian culture. They consistently picked the sentences that matched common stereotypes.

  • The Metaphor: Imagine a robot that has read every newspaper in India. It knows that "Caste" and "Religion" are huge topics in Indian news. So, when asked in English, it confidently picks the stereotypical answer.
  • The Catch: The robots were actually picking up on pan-Indian stereotypes (general ideas about India) rather than hyper-local Goan knowledge. For example, they knew stereotypes about "Hindus" or "Muslims" (which appear everywhere in India), but they were much worse at guessing stereotypes about very specific Goan groups like Tarvottis (seafarers) or Mundkars (tenants), which don't appear much in general Indian data.

In Konkani: The Robots Hit a Wall
When the researchers asked the same questions in Konkani, the robots suddenly became clueless. Their answers were basically random guesses (like flipping a coin).

  • The Metaphor: It's like asking a person who only speaks English to solve a math problem written in a language they've never seen. They don't say, "I know the answer but I'm choosing to be nice." They simply don't understand the language.
  • The Finding: The robots didn't have "no bias" in Konkani; they had no language skills in Konkani. They couldn't even tell if a sentence made sense or was nonsense.

3. The "Cultural Competence" Gap

The paper highlights a funny but serious gap:

  • General Multilingual Models: They are terrible at Konkani because they haven't seen enough of it in their training data.
  • Indian-Specific Models: These models are great at reading Konkani (they understand the grammar and words), but they still don't know Goan culture. When they read in Konkani, they don't show the stereotypes because their "Goan knowledge" is missing, not because they are morally superior.

4. Why This Matters

The authors used a creative analogy in their findings: Tokenization.
Think of words as LEGO bricks. If a robot sees the word "Tarvotti" (a specific Goan group), a good robot sees it as one solid brick. A bad robot sees it as a pile of broken, tiny pieces.

  • The researchers found that even when the robots could "read" the Konkani words (the LEGO bricks were whole), they still didn't know the cultural meaning behind them.
  • This proves that just because a robot can read a language, it doesn't mean it understands the people who speak it.

Summary

The paper concludes that:

  1. Bias is everywhere: AI models trained on English data have picked up on Indian stereotypes.
  2. Local knowledge is missing: These models don't know the tiny, specific details of Goan life (like specific job titles or local caste names).
  3. Language isn't enough: Just because a model can speak a low-resource language (like Konkani) doesn't mean it understands the culture. It might just be guessing.

The team released their test data (AmchiBias) so others can build better, more culturally aware robots that don't just guess, but actually understand the unique, complex identities of places like Goa.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →