Gender Disparities in LLM-Based Intimate Partner Violence Detection
This study demonstrates that large language models exhibit significant gender biases in detecting intimate partner violence, systematically underestimating abuse and perpetrator intent when the victim is male and the perpetrator is female, thereby reflecting gendered stereotypes from their training data.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have four very smart, well-read digital assistants (AI models) that have been trained on the entire internet. You want to test if they can spot when someone is being abused in a relationship. To do this, you take 475 real stories from people asking for advice online and play a game of "what if."
The Experiment: The Gender Swap
Think of each story as a movie script. In the original script, the victim and the abuser have specific genders (like a man and a woman). The researchers took these scripts and created four different "versions" of the same story for each one, simply by swapping the names and pronouns:
- Female Victim / Female Abuser
- Female Victim / Male Abuser
- Male Victim / Female Abuser
- Male Victim / Male Abuser
The actual words describing the bad behavior (the insults, the control, the threats) stayed exactly the same. Only the gender labels changed. Then, they asked the four AI models to read these scripts and answer questions like: "Is this abuse?" "Did the abuser mean to hurt them?" and "Is this a romantic relationship?"
The Big Discovery: The "Male Victim" Blind Spot
The results showed that the AI models aren't just looking at the actions; they are also looking at the genders of the people involved. It's as if the models have a hidden rulebook in their heads that says, "Abuse usually looks like a man hurting a woman."
When the story involved a male victim and a female abuser, the AI models were much less likely to say, "Yes, this is abuse."
- The "Invisible" Victim: If a man described being controlled or hurt by a woman, the AI often downplayed it. It was less likely to see the woman as having "bad intentions" or to label the situation as violence.
- The "Visible" Victim: Conversely, when the victim was a woman and the abuser was a man, the AI was much quicker to flag it as abuse.
The "Double Standard" Analogy
Imagine a security guard at a club (the AI) checking people's IDs.
- If a woman says, "My boyfriend is yelling at me," the guard immediately calls the police (flags it as abuse).
- If a man says, "My girlfriend is yelling at me," the guard shrugs and says, "Oh, that's probably just a misunderstanding," even if the words are identical.
The paper found that the AI models act like this guard. They seem to have absorbed the bias from the internet data they were trained on, which often treats male victims as less "real" or less in need of help.
Model Differences
Not all the AI models behaved the same way:
- Some models (like Llama) were very sensitive to these gender swaps, changing their answers drastically depending on who was who.
- Others (like GPT and Grok) were a bit more consistent but still showed a tendency to miss abuse when the victim was male.
- Interestingly, when the story involved two women (Female/Female), some models actually became more likely to spot the abuse, almost as if they were over-correcting because the "standard" male-victim scenario wasn't there.
The Bottom Line
The paper concludes that these AI tools are not neutral observers. They have learned to see the world through a gendered lens, where they expect abuse to happen in a specific way (man hurting woman). Because of this, if a man goes to an AI for help regarding a female partner, the AI might fail to recognize the danger, potentially leaving him without the support he needs. The authors warn that before we let these AIs help people in crisis, we need to fix these blind spots so they treat all victims equally.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.