← Latest papers
💬 NLP

IndoBias: A Dual Track Culturally Grounded Benchmark for LLMs Bias Evaluation in Indonesian Languages

This paper introduces IndoBias, a culturally grounded benchmark featuring dual evaluation tracks to assess LLM bias across Indonesian and three local languages, revealing that existing models exhibit significant representational unfairness and that pretraining on Common Crawl data or incorporating local languages can exacerbate these biases.

Original authors: Ikhlasul Akmal Hanif, Muhammad Falensi Azmi, Filbert Aurelian Tjiaranata, Eryawan Presma Yulianrifat, Fajri Koto

Published 2026-06-02
📖 4 min read☕ Coffee break read

Original authors: Ikhlasul Akmal Hanif, Muhammad Falensi Azmi, Filbert Aurelian Tjiaranata, Eryawan Presma Yulianrifat, Fajri Koto

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a giant, super-smart robot librarian who has read almost everything ever written on the internet. You ask this robot, "Who is the best kind of employee?" or "What are people from this village like?" The robot answers instantly, but sometimes, its answers are based on old, unfair, or exaggerated stories it heard while reading, rather than the truth. This is called bias.

For a long time, scientists have checked if this robot is biased, but they mostly only asked questions in English. They didn't check if the robot was fair to people in Indonesia, a country with over 1,300 different ethnic groups and 700 local languages. It's like checking a weather forecast only for London and assuming it's accurate for the entire planet.

This paper introduces IndoBias, a new "fairness test" designed specifically for Indonesia and its local languages (Javanese, Sundanese, and Makasar). Think of IndoBias as a two-part detective game to see how the robot thinks.

The Two Tracks of the Detective Game

The researchers built two different ways to test the robot:

1. The "Spot the Difference" Track (IndoBias-Pairs)
Imagine showing the robot two sentences side-by-side:

  • Sentence A (The Stereotype): "People from [Group X] are known for being [Lazy]."
  • Sentence B (The Counter-Story): "People from [Group X] are known for being [Hardworking]."

The robot has to pick which one "feels" more natural or likely to be true. If the robot consistently picks the negative stereotype, it's biased. The researchers created 544 pairs of these sentences in four languages. They found that the robot is very good at picking up on common stereotypes in Indonesian, but it gets even more biased when talking about religion and politics in local languages.

2. The "Open-Ended Story" Track (IndoBias-QA)
This track is like asking the robot to write a story or fill out a form about a specific person or group (like a specific tribe, a government office, or a university).

  • The Challenge: Some groups are famous and have clear stereotypes (like the Javanese). Others are smaller or less known (like the Korowai people).
  • The Finding: The robot treats these groups very differently. For famous groups, it might have a "standard" bias. But for smaller or marginalized groups, the robot often paints a much darker, less fair picture. It's as if the robot has a "spotlight" that shines brightly on some groups but leaves others in the dark, making them look worse than they are.

What Did They Discover?

The researchers ran some experiments to see why the robot acts this way:

  • The "Raw Internet" Problem: They found that if they train the robot on raw, unfiltered internet data (like Common Crawl), it learns more bias. It's like teaching a child by letting them read every comment section on the internet without a filter. In contrast, training on carefully edited sources like Wikipedia or news articles results in a slightly fairer robot.
  • The "Local Language" Surprise: You might think teaching the robot more local languages would make it fairer. However, the study found the opposite. When they added local languages (like Javanese and Sundanese) to the training mix, the robot actually became more biased. It seems the robot absorbed the real-world prejudices and stereotypes that exist in those specific local conversations.
  • Decoder vs. Encoder: The study compared different types of robot brains. The "Decoder" models (the ones that generate text and chat) were much more biased than the "Encoder" models (the ones that just understand text).

The Big Picture

The main takeaway is that bias isn't one-size-fits-all. A robot might be fair in English but unfair in Indonesian, or fair about one ethnic group but unfair about another.

The authors warn that if we don't test these robots in their specific cultural contexts, we risk building tools that accidentally hurt the very people they are supposed to help. They created this benchmark (IndoBias) so developers can check their robots against these specific cultural realities before releasing them to the public.

In short: You can't just translate an English fairness test to Indonesia and expect it to work. You need a custom-made test that understands the unique, complex, and diverse social fabric of the country. This paper provides that test.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →