Beyond Binary Classification: Detecting Fine-Grained Sexism in Social Media Videos
This paper introduces FineMuSe, a new Spanish multimodal dataset with fine-grained annotations and a hierarchical taxonomy, to advance sexism detection beyond binary classification by evaluating multimodal LLMs that show competitive performance with humans in identifying nuanced sexism while highlighting challenges in detecting co-occurring types conveyed through visual cues.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a robot how to spot rude or unfair behavior in a crowded, noisy video chat room.
For a long time, the robots we built were like very strict bouncers. They could only say two things: "You're in" (this is fine) or "You're out" (this is sexist). If a video was clearly hateful, the bouncer kicked it out. But if the sexism was subtle—like a joke that sounded funny but was actually mean, or a comment that sounded like a compliment but was actually a trap—the bouncer would let it slide. They were too binary; they only saw black and white, missing all the shades of gray.
This paper introduces a new, much smarter system called FineMuSe (think of it as a "Fine-Tuned Microscope for Sexism"). Here is how it works, broken down simply:
1. The New Dataset: A "Museum of Mean Moments"
The researchers collected 828 short videos from places like TikTok, YouTube Shorts, and BitChute. But instead of just labeling them "bad" or "good," they created a massive, detailed catalog.
Imagine a detective's board covered in red string. They didn't just say, "This video is sexist." They asked:
- Is it a Stereotype? (e.g., "Women are bad at math.")
- Is it Denial? (e.g., "There is no such thing as gender inequality anymore.")
- Is it Discrimination? (e.g., Attacking someone for being LGBTQ+.)
- Is it Objectification? (e.g., Treating a person like a piece of furniture to be looked at.)
- Is it a Joke? (Is the sexism hidden inside a punchline?)
They also labeled videos that fight back against sexism or report bad experiences. This creates a rich map of exactly how sexism shows up, not just that it shows up.
2. The Challenge: The "Visual Trap"
The researchers tested the world's smartest AI brains (Large Language Models) on this new map. They gave the AI two types of clues:
- The Script: Just the words spoken in the video.
- The Full Picture: The words, the tone of voice, and the actual video images.
The Result: The AI got pretty good at reading the script. If someone said, "Women should stay home," the AI knew that was sexist.
The Problem: The AI struggled when the sexism was hidden in the visuals.
Imagine a video where a man says, "I love my wife," which sounds nice. But the video shows him treating her like a servant, carrying her bags while she sits on a throne. The AI read the nice words but missed the mean picture. It's like a person who only listens to the lyrics of a song but doesn't see the angry face of the singer. The AI needs to learn to "watch" the video, not just "read" the transcript.
3. The "Human vs. Robot" Showdown
The researchers also asked: "Can the AI explain why it flagged a video?"
They compared the AI's explanations to those written by human experts.
- The Verdict: The AI is surprisingly good at writing explanations! It can say, "This is sexist because it uses a joke to hide a stereotype," almost as well as a human.
- The Catch: Sometimes the AI gets the reason right but misses the visual part of the reason. It's like a detective who correctly identifies the crime but forgot to look at the fingerprint on the window.
4. Why This Matters
Think of online moderation like trying to filter water.
- Old Way: You have a big net that catches only the biggest rocks (obvious hate speech). The tiny, invisible sand (subtle sexism) slips right through, making the water still dirty.
- New Way (FineMuSe): You have a high-tech filter that catches the rocks, the sand, and even the microscopic algae. It understands that sexism isn't just one thing; it's a chameleon that changes colors depending on the platform (TikTok vs. YouTube) and the language (Spanish dialects from Mexico vs. Spain).
The Bottom Line
This paper is a giant leap forward. It tells us that while AI is getting very smart at spotting obvious sexism, it still needs to learn how to "read the room" visually. It's not enough to just hear the words; the AI needs to understand the body language, the context, and the hidden jokes.
By creating this detailed "museum" of examples, the researchers are giving future AI a better textbook to study, hoping to build tools that can finally catch the subtle, sneaky forms of unfairness that have been hiding in plain sight.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.