Sexualised synthetic personas encode and amplify gendered power asymmetries through voice
Drawing on a Feminist HCI perspective, this study demonstrates that commercial AI voice systems reinforce gendered power asymmetries by encoding female-coded voices with sexualized and submissive traits while associating male-coded voices with dominance and positive attributes.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a digital toy store where you can buy a "voice" instead of a teddy bear. You can pick a voice that sounds like a "Southern Gentleman" or a "Parisian Temptress." You can even tell the voice to giggle, sigh, or sound "flirty." This is what a popular company called ElevenLabs offers.
This research paper is like a group of detectives (the researchers) walking into that store, buying a bunch of these voices, and then asking a diverse group of regular people: "What do these voices actually make you feel?"
Here is what they found, broken down simply:
1. The Setup: The "Voice Menu"
The researchers looked at the most popular voices on the site. They noticed the menu was split into two main sections:
- The "Serious" Section: Voices labeled "Educational," "News Anchor," or "Tech Expert."
- The "Sexy" Section: Voices labeled "Flirty," "Temptress," or "Seductive."
They picked the top male and female voices from both sections. Then, they made a clever switch: they took the "sexy" voices and made them read boring, neutral text (like a science passage). They also took the "serious" voices and made them read the same boring text. This was like testing if a "sexy" voice sounds sexy even when it's talking about math, or if a "serious" voice sounds serious even when the words are neutral.
2. The Experiment: The "Word Game"
They played these audio clips for about 120 people in the US and Canada. The listeners had a big list of 36 words (like charming, creepy, dominant, shy, seductive) and had to pick the three that best described what they heard.
3. The Big Discovery: The "Gender Script"
The results showed that the voices weren't just different; they were following a very old, very strict script.
The Male Voices (The "Boss" Script):
Whether they were reading a sexy script or a boring science script, the male voices were mostly described as dominant, confident, and positive. Even when they were trying to sound "flirty," people still heard a guy in charge.- Analogy: Think of a male voice like a sturdy oak tree. Whether it's in a storm or a sunny day, it still feels solid, strong, and reliable.
The Female Voices (The "Submissive" Script):
The female voices were almost always described as submissive, shy, and highly sexualized. Even when they were reading the boring science text, people still heard them as "flirty" or "seductive."- Analogy: Think of a female voice like a rubber duck. No matter what you put it in (a bathtub or a kitchen sink), it's still expected to squeak, float, and be "cute." The researchers found that the female voices had extra "fluff" added to them—like breathy sighs and giggles—that made them sound submissive, regardless of what they were actually saying.
4. The "Content" vs. The "Tone"
The researchers wanted to know: Is it the words that make the voice sound sexy, or is it the sound itself?
- For Men: The words mattered a lot. If a male voice read a sexy script, people called him "seductive." If he read a boring science script, people called him "serious." The words changed the perception.
- For Women: The words barely mattered. Whether a female voice read a sexy script or a boring science script, people still called her "seductive," "intimate," or "sensual."
- The Takeaway: The "sexiness" was baked into the voice's tone and breathing, not just the words. The female voice was "programmed" to sound like a fantasy, even when it was just reading a textbook.
5. What the Listeners Said (The "Real Talk")
The researchers also asked people to write down their honest thoughts.
- About the Men: Some people thought the "sexy" male voices sounded "threatening" or "predatory," while others thought the "serious" ones were just "boring" or "flat."
- About the Women: Many people felt the "sexy" female voices were cringe-worthy. They used words like "forced," "trying too hard," and "annoying." One listener said it sounded like "a teenage boy's idea of what sexy sounds like."
- The Metaphor: It was like watching a bad actor trying to play a seductress. It felt like a caricature—an exaggerated, fake version of femininity that didn't feel real or human.
6. The Conclusion: Why This Matters
The paper argues that these AI voices are like digital mannequins. They aren't showing us the full, messy, diverse reality of how real men and women speak. Instead, they are reinforcing old stereotypes:
- Men are the bosses (dominant, positive).
- Women are the objects (submissive, sexualized).
The researchers warn that even though these AI voices are marketed as "creative tools" for making video games or movies, they are actually copying and amplifying harmful biases. They are taking the "male gaze" (how men imagine women should sound) and turning it into a product that anyone can buy.
In short: The paper found that while AI voice technology could be a tool for freedom and diversity, right now, it's mostly acting like a mirror that only reflects outdated, stereotypical ideas about gender, making women sound like submissive fantasies and men sound like dominant bosses.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.