Bielik Guard: Efficient Polish Language Safety Classifiers for LLM Content Moderation
The paper introduces Bielik Guard, a family of efficient, compact Polish language safety classifiers fine-tuned on a community-annotated dataset that effectively categorize content into five safety domains while significantly outperforming existing models in precision and false positive rates.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you've built a massive, incredibly smart library where books can talk back to you. These are Large Language Models (LLMs)—AI chatbots that can write stories, solve math problems, and chat about anything. But, just like a public library, you need a librarian to make sure no one is shouting obscenities, planning crimes, or hurting themselves inside the building.
The problem? Most of the best librarians in the world only speak English. If you try to use an English librarian to watch over a Polish conversation, they might miss subtle jokes, misunderstand local slang, or get too scared of harmless words, blocking good conversations by mistake.
Enter Bielik Guard.
The Story of Bielik Guard
The authors of this paper created a team of specialized Polish librarians called Bielik Guard. The name comes from "Bielik," a majestic eagle, and the code name "Sójka," which means a Jay (a bird known for being loud, watchful, and protective).
Their goal was simple: Build a safety system that speaks Polish fluently, understands Polish culture, and is small enough to run on a regular computer without needing a supercomputer.
How They Built It: The "Community Crowd"
Instead of hiring a few expensive experts to write the rulebook, the team asked 1,500 regular Polish people to help. They set up a simple online survey where volunteers read short snippets of text and voted on whether they were "safe" or "unsafe."
Think of it like a neighborhood watch. Instead of one person deciding what's dangerous, the whole neighborhood votes. If 6 out of 10 neighbors think a joke is too mean, the system learns that it's a problem. If only 2 neighbors think it's bad, the system learns it's probably fine. This approach captures the nuance of the Polish language—things like sarcasm, slang, and cultural context that a foreign AI would miss.
They collected about 6,885 unique texts and over 60,000 votes to train their system.
The Two Models: The "Pocket Watch" and the "Grandfather Clock"
The team built two versions of their safety guard:
Bielik Guard 0.1B (The Pocket Watch):
- Size: Tiny (124 million parameters). It's like a smartwatch.
- Speed: Super fast. It can check text almost instantly.
- Job: Great for apps where speed and battery life matter.
- Performance: Surprisingly, this tiny model is smarter at spotting Polish problems than much larger models from other countries.
Bielik Guard 0.5B (The Grandfather Clock):
- Size: Medium (443 million parameters). It's like a sturdy wall clock.
- Speed: Slightly slower, but still very fast.
- Job: The heavy lifter. It catches more subtle dangers and handles tricky, messy text better.
What They Watch For
The Polish librarians are trained to spot five specific types of trouble:
- Hate/Aggression: Bullying or attacking groups of people.
- Vulgarities: Swearing and dirty language.
- Sexual Content: Explicit descriptions or requests for erotica.
- Crime: Instructions on how to make drugs or commit fraud.
- Self-Harm: Anything encouraging suicide or eating disorders.
Note: They intentionally ignore things like "fake news" or "copyright" because those require different kinds of thinking. They focus on immediate safety.
The Big Test: Why They Win
The team tested their librarians against the "giants" of the industry (like Llama Guard and QwenGuard), which are huge models designed to speak 20+ languages.
Here is the result, explained simply:
- The Giants: Because they try to speak every language, they are often confused by Polish. They shout "STOP!" too often, blocking innocent conversations (high False Positives). Imagine a guard who stops everyone because they look slightly suspicious.
- Bielik Guard: Because it was trained specifically on Polish data, it knows exactly what to look for.
- Precision: When Bielik Guard says, "This is dangerous," it is 77% right.
- The Giants: When the English-trained giants say, "This is dangerous," they are often wrong (only ~13% right on Polish text).
- The Result: Bielik Guard lets normal people talk freely without getting blocked, while still catching the real bad actors. It's like having a guard who knows the difference between a playful shove and a real fight, whereas the foreign guard thinks a playful shove is a fight.
A Special Feature: Helping, Not Just Blocking
Most safety systems just say "NO" and block the message. But for the Self-Harm category, Bielik Guard is designed to be different.
If someone writes something sad or dangerous, instead of just silencing them, the system is designed to help. It can be set up to provide a phone number for a crisis helpline (like "Telefon Zaufania" in Poland). It treats the user like a person in need of support, not just a rule-breaker.
The Bottom Line
Bielik Guard proves that you don't need a giant, expensive AI to keep a specific language safe. By using local data, community votes, and smart, small models, you can build a safety system that is:
- More accurate for Polish speakers.
- Less annoying (it doesn't block innocent people).
- Cheaper and faster to run.
It's a reminder that for language safety, knowing the culture is often more important than having a bigger brain.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.