Vision-Language Models Suppress Female Representations Under Ambiguous Input
This paper reveals that while vision-language models often internally encode female associations for ambiguous images, their outputs systematically default to male representations due to an asymmetric filtering mechanism that amplifies male signals and suppresses female ones during generation.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Idea: The "Polite" Robot vs. The "Biased" Brain
Imagine you have a very polite robot assistant who has been trained to be fair and not make assumptions about people. If you show it a picture of a person clearly wearing a dress, it says, "That's a woman." If the person is clearly in a suit, it says, "That's a man." It passes the "politeness test" perfectly.
But this paper asks a tricky question: What happens when the robot can't see the person's face or gender clearly? (For example, a construction worker seen from behind in a helmet, or a nurse seen from the back).
The researchers found that while the robot's mouth (its final answer) stays polite and neutral, its brain (its internal thinking process) is actually making a very specific, biased guess: It assumes the person is a man.
Even for jobs that are statistically mostly done by women (like babysitters or florists), if the robot has to guess, it guesses "man." The scary part? The robot knows the job is usually female, but it overrides that knowledge and forces a "male" answer anyway.
The Tool: "LALS" (The X-Ray for Thoughts)
How do we know what the robot is thinking if it's not saying it? The authors invented a tool called LALS (Latent Association Leaning Score).
Think of LALS as an X-ray for the robot's brain.
- Normally, we only see the robot's final sentence (the output).
- LALS lets us peek inside the robot's "neurons" at every single step of its thinking process.
- It translates the robot's internal electrical signals into words, allowing us to see if, at a specific moment, the robot is thinking "female," "male," or "neutral."
The Three Types of Jobs
When the researchers tested 15 different jobs with "faceless" images, they found three distinct patterns:
The "Agree" Men (e.g., Firefighters):
- Inside the brain: The robot thinks "Man."
- The answer: It says "Man."
- Verdict: No surprise here. The bias is consistent.
The "Agree" Women (e.g., Makeup Artists):
- Inside the brain: The robot thinks "Woman."
- The answer: It says "Woman."
- Verdict: Also consistent.
The "Divergence" Zone (e.g., Florists, Nurses, Babysitters):
- Inside the brain: The robot actually thinks "Woman." (It correctly associates the job with women).
- The answer: It says "Man."
- Verdict: This is the hidden bias. The robot's internal thought says "female," but something in its final step forces it to flip the switch and say "male."
The "Asymmetric Filter": A One-Way Street
The researchers discovered why this flip happens. They traced the robot's thoughts layer by layer (like checking the floors of a skyscraper).
- The Male Signal: Imagine a strong, loud voice shouting "MAN!" This voice starts at the bottom floor and gets louder as it goes up. By the time it reaches the top floor (where the answer is spoken), it is very strong.
- The Female Signal: Imagine a different voice shouting "WOMAN!" This voice starts strong, gets even louder in the middle of the building, but then hits a soundproof wall on the top floors. By the time it reaches the top, the voice has been silenced or turned into a whisper.
The robot has an asymmetric filter. It lets the "Male" idea pass all the way through, but it actively suppresses the "Female" idea right before it speaks.
The "Pink vs. Blue" Clue
To prove the robot was actually looking at the picture and not just guessing randomly, the researchers played a trick with colors.
- They showed a picture of a nurse in blue scrubs. The robot's internal "Male" signal was strong.
- They showed the exact same picture but changed the scrubs to pink.
- Result: The robot's internal "Male" signal dropped significantly, and the "Female" signal grew stronger.
This proves the robot isn't just blindly guessing "Man" for everything. It is reacting to cultural cues. It has learned that in our world, pink often means "female" and blue often means "male." When the visual cue (pink) is strong, the robot's internal bias shifts. But without that strong cue, it defaults to "Man."
The Root Cause: It's Not Just "Training"
The researchers checked if this bias was taught to the robot to be "safe" (a process called alignment). They compared the robot before it was taught to be polite and after.
- The Finding: The bias was there before the robot was taught to be polite.
- The Conclusion: The "polite training" didn't fix the bias; it just taught the robot to hide it. The robot learned to say "I don't know" or "Man" to be safe, but its internal brain still carries the old, biased associations.
Summary
The paper reveals that Vision-Language Models have a hidden double life.
- Publicly: They are neutral and careful.
- Privately: When they can't see clearly, they default to assuming people are men, even when their own internal logic suggests the person is a woman.
The "Male Default" isn't a glitch; it's a deep-seated habit in the robot's brain that survives even when the robot is told to be fair. The robot knows the truth (that the person might be a woman), but it filters that truth out before it speaks.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.