← Latest papers
💬 NLP

From Preferences to Prejudice: The Role of Alignment Tuning in Shaping Social Bias in Video Diffusion Models

This paper introduces **VideoBiasEval**, a diagnostic framework that reveals how alignment tuning—while improving video quality—unintentionally amplifies and stabilizes social biases by propagating stereotypes from human preference datasets through reward models into video diffusion models.

Original authors: Zefan Cai, Haoyi Qiu, Haozhe Zhao, Ke Wan, Jiachen Li, Jiuxiang Gu, Wen Xiao, Nanyun Peng, Junjie Hu

Published 2026-02-12
📖 3 min read☕ Coffee break read

Original authors: Zefan Cai, Haoyi Qiu, Haozhe Zhao, Ke Wan, Jiachen Li, Jiuxiang Gu, Wen Xiao, Nanyun Peng, Junjie Hu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a robot how to cook. At first, the robot just watches random videos of people in kitchens. It’s a bit messy, and sometimes it shows a person chopping an onion, but other times it shows a person dancing. It’s not "bad," but it’s not very consistent.

To make the robot better, you decide to give it a "Preference Guide." You show it thousands of videos and say, "People liked these videos more because they look professional and clean." This process is called Alignment Tuning.

This paper, "From Preferences to Prejudice," discovers a hidden problem: When we teach the robot to be "better" based on what humans "prefer," we accidentally teach it to be biased.

Here is the breakdown of their discovery using three simple analogies:

1. The "Echo Chamber" Effect (How Bias Travels)

Think of bias like a whisper in a crowded room.

  • The Whisper (Human Data): When humans rate images, they often unconsciously prefer certain looks (for example, favoring certain ethnicities or genders).
  • The Megaphone (Reward Models): When we train a "Reward Model" (the robot's guide) on those human ratings, the model doesn't just repeat the whisper; it turns it into a shout. It learns that "High Quality = This specific type of person."
  • The Broadcast (Video Models): Finally, when the video generator tries to please the "Megaphone," it stops being diverse entirely. It starts producing the same "preferred" types of people over and over again.

The paper shows that alignment doesn't just fix the video quality; it "locks in" the stereotypes.

2. The "Polished Stereotype" (The Danger of Smoothness)

This is the most chilling finding in the paper. Usually, when an AI makes a mistake, it looks like a glitch—a weirdly shaped hand or a flickering face. You can see it's an error.

However, the researchers found that Alignment Tuning acts like a professional makeup artist for stereotypes.

Before alignment, a model might be biased, but it’s "messy" and inconsistent. After alignment, the model becomes much better at making smooth, high-quality, beautiful videos. But because it is now so "good" at following the biased guide, it produces perfectly polished, high-definition stereotypes. It’s no longer a glitch; it’s a smooth, convincing, and "professional-looking" prejudice. This makes the bias much harder to spot and more dangerous because it looks "correct."

3. The "Volume Knob" (Can we fix it?)

The researchers asked: "If we can turn the bias up, can we turn it down?"

They treated bias like a Volume Knob. They created a "Man-Preferred" guide and a "Woman-Preferred" guide.

  • If they gave the robot the "Man-Preferred" guide, the volume of male representations went up.
  • If they gave it the "Woman-Preferred" guide, they could actually "counter-steer" the robot and force it to show more women.

The takeaway: They proved that bias in AI isn't a permanent broken part of the machine; it's a setting. We can use this "knob" to intentionally design AI that is more diverse and fair, rather than just letting it follow the accidental prejudices of the internet.

Summary in one sentence:

When we teach AI to follow human "preferences" to make videos look better, we are accidentally teaching it to turn messy human prejudices into smooth, high-definition stereotypes—but we also learn that we can use those same tools to steer the AI toward fairness.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →