BAFIS: Dataset + Framework to assess occupational Bias and Human Preference in modern Text-to-image Models
This paper introduces BAFIS, a dataset and framework comprising 21,140 multilingual images and human preference feedback, to evaluate and reveal systematic occupational biases in five leading text-to-image models, demonstrating the critical need to incorporate human judgment alongside established metrics for developing fairer AI systems.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a magical painting machine. You type in a description like "a doctor," and it instantly paints a picture. For a long time, people thought these machines were just neutral tools, like a camera. But this paper argues that these machines are more like biased mirrors—they don't just show reality; they show a distorted version of it based on what they've learned from the internet.
The authors of this paper, Thomas, Adrian, and Biying, wanted to test five of the most popular "painting machines" (Midjourney, Stable Diffusion, DALL-E 3, Playground, and FLUX) to see how they handle jobs and people. They built a special testing ground called BAFIS (Battle-Arena for Fair Image Synthesis).
Here is a breakdown of their findings using simple analogies:
1. The Arena: BAFIS
Think of BAFIS as a blind taste-test competition.
- The Setup: Instead of just letting a computer grade the pictures, they put two different machines in a "battle."
- The Contest: A user is shown a prompt (e.g., "a nurse") and two anonymous images generated by two different machines.
- The Vote: The user votes on three things:
- Bias: Does the picture look like a realistic mix of people, or does it only show one type?
- Quality: Is the picture pretty and clear?
- Alignment: Does the picture actually match the words you typed?
- The Score: They used a ranking system (like chess ratings) to see which machine won the most votes.
2. The Dataset: A Giant Library of Jobs
To make sure the test was fair, they didn't just ask for "a doctor." They created a massive library of 21,140 pictures based on 1,057 different job descriptions.
- They used two languages: English and German.
- They tried different ways of asking:
- Direct: "A photo of a male accountant."
- Indirect: "A photo of a person who manages finances."
- Groups: "A photo of a group of accountants."
- They compared the results against real-world statistics from the German government (like checking a recipe against the actual ingredients list) to see if the machines were lying about who holds these jobs.
3. The Findings: What the Machines Got Wrong
The Gender Problem (The "Construction vs. Nursing" Split)
The machines have strong stereotypes about gender.
- Construction & Engineering: When asked to draw these jobs, the machines almost always drew men. It was like they thought only men could build things.
- Healthcare & Office Work: When asked to draw nurses or office workers, the machines leaned heavily toward women.
- The Surprise: One machine, DALL-E 3, was the most balanced. It drew men and women almost equally, like a fair coin flip. Another machine, FLUX, actually drew more women than men for some jobs, which is a new kind of bias.
The Ethnicity Problem (The "White Wall")
This was the biggest issue. No matter what job you asked for, the machines overwhelmingly drew white faces.
- It was like the machines had a default setting that said, "If you don't specify a race, make them white."
- This bias was even stronger in German than in English. When the prompt was in German, the machines drew even fewer non-white faces.
- DALL-E 3 was the only machine that managed to keep the percentage of white faces below 50% in both languages, making it the most diverse.
The "Human vs. Robot" Scorecard
Here is where it gets tricky. The paper found that computer metrics (math formulas that grade image quality) often disagreed with human feelings.
- Image Quality: A machine called FLUX was rated by humans as the most beautiful and realistic. However, the computer math (FID score) said DALL-E 3 was the best.
- The Takeaway: The math formulas are like a robot trying to judge a painting; they measure pixels but miss the "soul" or the feeling that humans get when looking at an image. Humans preferred FLUX, even if the math said otherwise.
The Language Twist
- DALL-E 3 has a secret weapon: it rewrites your prompt. If you type a German prompt, it secretly translates and tweaks it into English before painting. This helped it understand the request better and get higher scores for "alignment" (matching the words) in German battles.
- Stable Diffusion 3 was the champion for English prompts, but it struggled with German.
4. The Conclusion
The paper concludes that these AI painting machines are not neutral. They carry the biases of the data they were trained on.
- They over-represent men in "hard" jobs and women in "care" jobs.
- They almost always default to white faces.
- Computer tests aren't enough. We need to ask real humans what they think because the math doesn't always capture what feels "fair" or "real" to us.
The authors suggest that to fix this, developers need to listen to human feedback more closely and check their machines against real-world statistics, not just computer formulas. They built BAFIS so that in the future, we can keep holding these machines accountable in a "battle arena" to ensure they become fairer.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.