Investigating Bias in Bulgarian in the Context of Large Language Models
This paper presents a manually annotated Bulgarian dataset to evaluate bias detection, revealing high agreement among human annotators but significant divergence between humans and large language models, while demonstrating that Bulgarian-language prompting improves model alignment with human judgments.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are teaching a super-smart robot how to understand human stories. You give it millions of books, articles, and websites to read, hoping it learns not just grammar, but also how to be fair, kind, and neutral. This is the world of Large Language Models (LLMs), the AI brains behind chatbots and search engines. But here's the catch: robots learn by mimicking what they see. If the books they read contain old-fashioned stereotypes or unfair opinions about certain groups of people, the robot might accidentally learn those biases too. It's like a student who only reads history books written by one side of a war; they might think that side is always right. Scientists are worried because if these robots start making decisions about jobs, loans, or news, they could accidentally treat people unfairly. The big question is: Can we teach these robots to spot unfairness in text, and do they see it the same way humans do?
This paper dives into that question, but with a twist: it looks at the Bulgarian language. Most AI research focuses on English, which is like studying only one flavor of ice cream and assuming you understand the whole freezer. The researchers wanted to see if the same rules apply to a language with fewer digital resources and a very different culture. They built a special "test kit" using 3,177 sentences from Bulgarian Wikipedia. Two human experts read these sentences and graded them on a scale of 0 to 5 to see how biased they were, looking for unfairness related to gender, religion, race, appearance, and disability. Then, they asked three different AI models (Gemma 4, BgGPT-Gemma-3, and Qwen 3) to grade the exact same sentences.
The results were a bit of a shock, like asking a human and a robot to judge a magic trick and getting completely different answers. The two humans agreed with each other almost all the time (only 4.9% of the time they disagreed on whether a sentence was biased). However, when the AI models tried to do the same job, they were much less reliable. The robots disagreed with the humans 24% to 34% of the time. Even worse, when the humans said a sentence was biased, the robots often didn't see it, and when the robots did flag something as biased, the humans often thought it was fine.
The paper suggests that the robots aren't just "wrong"; they are playing by a different rulebook. The humans were very strict about how bad the bias was (giving high scores for strong bias), but the robots were much more cautious, giving very low scores even when humans saw strong unfairness. It's as if the humans are shouting, "That's terrible!" while the robots are whispering, "Maybe a little rude?"
There was one interesting discovery, though. One of the AI models, BgGPT-Gemma-3, was built specifically for Bulgaria. When the researchers asked it questions in English, it did a poor job. But when they switched the instructions to Bulgarian, the robot suddenly started seeing bias more like a human did. It's as if speaking the robot's native language woke up a different part of its brain, helping it understand the cultural context better.
The study concludes that we can't just trust robots to fix bias in low-resource languages like Bulgarian yet. They are currently too different from human judgment. The author suggests that instead of relying on a single AI to do the grading, we might need to use a team of different AIs and then have humans check their work. They also emphasize that we need to build better dictionaries and tools specifically for Bulgarian culture, because the robots are currently missing the subtle, cultural clues that humans pick up on instantly. Until then, the robots are still learning how to be truly fair.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.