Learning Sexism Detection Using Multi-Agent Perspectivist Preference Optimization
This paper introduces the Multi-Agent Perspectivist Preference Optimization (MAP-PO) framework, which addresses annotator disagreement in sexism detection by clustering labeling behaviors and fine-tuning specialized language model agents that are coordinated through preference optimization to preserve diverse perspectives while maintaining alignment with majority labels.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a robot how to spot a mean joke. You ask ten different people to read a joke and decide if it's sexist. Some say "Yes, that's terrible," while others say "No, that's just a joke." In the old days of computer science, if the majority said "Yes," the computer would just ignore the "No" votes and call it a fact. But what if those "No" votes aren't mistakes? What if they are just different people seeing the world through different lenses? This is the heart of a field called "perspectivist" science. It argues that when humans disagree, that disagreement isn't noise to be deleted; it's a signal to be studied. The big question is: how do you build a computer that doesn't just pick a winner, but actually understands why different people see things differently?
This paper introduces a clever new system called MAP-PO (Multi-Agent Perspectivist Preference Optimization) to solve exactly that problem. Instead of trying to force one robot to be the "perfect" judge, the researchers built a team of three specialized robot agents. Here's how they did it: First, they looked at the people who labeled the data (the humans) and didn't group them by age or gender. Instead, they grouped them by how they voted. Some humans were very strict and said "Yes" often; others were more lenient. The researchers found that these "voting styles" were the real key to understanding the data, not the humans' backgrounds.
Next, they trained one robot agent to mimic each of these three voting styles. But here was the tricky part: if they just told each robot, "Only listen to your own group," the robots went crazy. They became extreme caricatures, with one robot saying "Yes" to everything and another saying "No" to everything, losing touch with the real humans they were supposed to represent. The paper found that to fix this, the robots needed a "team signal." They had to be rewarded not just for being true to their own style, but also for working together to get the right overall answer. When they added this team reward, the robots stayed balanced. They learned to be distinct personalities without losing their minds. In the end, this team of three robots was better at spotting sexism than any single robot or any old-fashioned computer program, proving that keeping different perspectives alive makes for a smarter, more honest system.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.