ConGA: Guidelines for Contextual Gender Annotation. A Framework for Annotating Gender in Machine Translation
This paper introduces the Contextual Gender Annotation (ConGA) framework, a linguistically grounded guideline for word-level gender annotation that addresses gender bias in machine translation from English to Italian by creating a gold-standard dataset to evaluate and improve the performance of current MT systems.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are a translator trying to explain a story from English to Italian.
In English, the language is a bit like a chameleon when it comes to gender. You can say, "The doctor is here," and the word "doctor" doesn't tell you if the person is a man or a woman. It's neutral. You can also say, "They are happy," and "they" could be anyone.
But in Italian, the language is like a strict dress code. You must pick a gender. You can't just say "doctor"; you have to say il dottore (male doctor) or la dottoressa (female doctor). You can't just say "happy"; you have to choose contento (male happy) or contenta (female happy).
The Problem: The "Default" Setting
The problem is that when computers (Machine Translation systems) try to translate from English to Italian, they often get stuck. Since the English source doesn't give them a clue, the computer panics and picks the "safe" option: the male version.
It's like a waiter who, when asked "What would you like to drink?" and the customer says "Something neutral," automatically pours a beer because "most people drink beer," even if the customer might have wanted juice. Over time, this makes the world in the computer's mind look like a place where everyone is a man, reinforcing old stereotypes (like thinking all nurses are women and all engineers are men).
The Solution: ConGA (The Detective's Notebook)
The authors of this paper created a new set of rules called ConGA (Contextual Gender Annotation). Think of ConGA as a detective's notebook for translators.
Instead of just guessing, the researchers went through thousands of sentences and put "sticky notes" on the words to mark exactly what the gender situation was.
The English Side (The Clues): They looked at the English sentence and asked: "Do we know the gender?"
- M (Masculine): Yes, it's a man.
- F (Feminine): Yes, it's a woman.
- A (Ambiguous): We don't know. It could be either. (This is the tricky part).
The Italian Side (The Action): They looked at the Italian translation and asked: "Did the computer make a choice?"
- If the English was Ambiguous (A), but the Italian translation forced a male or female ending, the detective notebook flags it.
- If the computer guessed "Male" 9 times out of 10 when it didn't know, the notebook records that as a bias.
The Experiment: Testing the Computers
The researchers took two powerful AI translators (one called TowerLLM and another called mBART) and asked them to translate these sentences. Then, they used their "detective notebook" to grade the work.
What they found:
- The "Male Default" is real: When the AI didn't know the gender, it almost always guessed "Male." It was like a coin that was weighted on one side.
- Missing the Females: The AI was much better at getting the "Male" guesses right when it did know the gender, but it often missed the "Female" clues entirely, turning a female nurse into a male nurse by accident.
- The Bias Score: They calculated a score (like a report card) showing that while the AI is getting smarter, it still has a heavy "male bias" habit that it can't break.
Why This Matters
This isn't just about grammar; it's about fairness.
If AI systems keep translating "The doctor" as "He" and "The nurse" as "She" (even when the English doesn't say so), they are teaching the world that men are the default humans and women are the exceptions.
ConGA is the tool that helps us catch this. It gives us a way to measure exactly how biased a computer is, so we can fix it. It's like putting a speed camera on a car that always speeds; once we can see the speed, we can finally build a system that drives fairly.
The Takeaway
This paper gives us a rulebook and a scorecard to stop computers from assuming everyone is a man just because they don't know better. By teaching the AI to recognize when it doesn't know the gender, and to leave it neutral (or ask for help) instead of guessing "Male," we can build a future where technology reflects the real, diverse world we live in.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.