Robust Bias Evaluation with FilBBQ: A Filipino Bias Benchmark for Question-Answering Language Models
This paper introduces FilBBQ, a culturally adapted Filipino bias benchmark comprising over 10,000 prompts that, when applied with a robust multi-seed evaluation protocol, reveals significant sexist and homophobic biases in Filipino language models regarding emotion, domesticity, queer interests, and polygamy.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a very smart, well-read robot that can speak many languages. You ask it questions, and it answers them like a human. But there's a catch: because this robot learned from the internet, it might have picked up some of society's unfair stereotypes, like "women are bad at math" or "gay men are obsessed with fashion."
This paper is about building a specialized test to see if this robot has these biases, but specifically for the Philippines and the Filipino language.
Here is the story of their work, broken down into simple concepts:
1. The Problem: The "One-Size-Fits-All" Test Didn't Fit
Scientists already had a famous test called BBQ (Bias Benchmark for Question-Answering). Think of BBQ like a standardized driving test. It's great for checking if a car (or a robot) knows the rules of the road in the US or Europe.
But the Philippines is a different country with different roads, different traffic signs, and different driving habits.
- The Gap: The existing tests were mostly in English or other major languages. They didn't capture the unique stereotypes found in Filipino culture. For example, in the US, a stereotype might be about "denim overalls," but in the hot Philippines, that's not a thing. In the Philippines, a specific type of local helper (yaya) or a local term for a queer man (bakla) carries different cultural weight.
- The Flaw in Testing: Previous tests asked the robot a question once and took the answer as the final truth. But robots are like nervous actors; if you ask them the same question twice in a row, they might give two different answers. Relying on just one answer is like judging a basketball player's skill based on a single shot.
2. The Solution: Building "FilBBQ" (The Filipino BBQ)
The researchers built a new test called FilBBQ. They didn't just translate the American test; they rebuilt it from the ground up to fit Filipino culture.
Think of this process like cooking a local dish:
- Step 1: Sorting the Ingredients. They looked at the original American test and threw away ingredients that didn't make sense in the Philippines (like stereotypes about sports that aren't popular there).
- Step 2: Localizing the Flavor. They translated the remaining questions but changed the names and situations. Instead of "Donna and Jermaine," they used popular Filipino names. Instead of "babysitter," they used "yaya" (nanny).
- Step 3: Adding New Spices. They added brand new questions about stereotypes unique to the Philippines, like the idea that "tomboys" (masculine-presenting women) are good with cars, or that queer men are obsessed with fashion and gossip.
- The Result: They created over 10,000 questions (prompts) to test the robot.
3. The New Rule: Don't Trust a Single Shot
The most important part of this paper is how they tested the robots.
- Old Way: Ask the robot a question once. If it says "Women are emotional," you mark it as biased.
- New Way (The Robust Protocol): The researchers asked the robot the same question 50 times (using different random seeds, like rolling dice to change the robot's mood).
- Why? Because robots are unstable. Sometimes they might say "Women are emotional," and other times "Men are emotional," or "I don't know." By averaging the results of 50 tries, they got a true picture of the robot's personality, rather than just catching it on a bad day.
4. What They Found (The Taste Test Results)
When they ran this new, rigorous test on three different Filipino-speaking robots, they found some interesting things:
- The "Emotional" Stereotype: The robots consistently thought that if a situation was vague, the person being emotional was likely a woman.
- The "Domestic" Stereotype: The robots strongly associated women with being nurses, homemakers, and taking care of the family, while associating men with being doctors or breadwinners.
- The "Queer" Stereotype: The robots often assumed that gay men were interested in fashion, design, and gossip, or that they struggled with monogamy (being faithful).
- The "Big Data" Surprise: The robot that had read the most Filipino text (the biggest brain) actually showed the strongest biases. It seems that the more a robot reads about the world, the more it absorbs the world's unfair stereotypes, too.
5. Why This Matters
This paper is like a mirror held up to technology.
- It shows us that if we want AI to be fair in the Philippines, we can't just use American tests. We need to understand local culture.
- It teaches us that we can't trust a robot's answer if we only ask it once. We need to ask it many times to see the truth.
- It gives us a tool (FilBBQ) to check if future robots are being fair to Filipino women and the LGBTQ+ community.
In a nutshell: The researchers built a culturally specific "lie detector" for robots in the Philippines and proved that to get an honest answer, you have to ask the same question over and over again. They found that even our smartest robots still carry the heavy baggage of old-fashioned stereotypes.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.