← Latest papers
💻 computer science

SLAyiNG: A Diverse and Community-validated Dataset of Queer Slang

This paper introduces "SLAyiNG," the first community-validated dataset of over 500 English queer slang terms across 20 subcommunities, which reveals that while many language models lack representation bias against queer identities, they still struggle to process queer vernacular—particularly slang from African-American and Latine communities—highlighting the need for diverse linguistic resources to improve NLP systems.

Original authors: Leonor Veloso, Lea Hirlimann, Lucija Mihić Zidar, Philipp Wicke, Valentin Hofmann, Hinrich Schütze

Published 2026-08-14
📖 3 min read☕ Coffee break read

Original authors: Leonor Veloso, Lea Hirlimann, Lucija Mihić Zidar, Philipp Wicke, Valentin Hofmann, Hinrich Schütze

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a world where computers are like giant, voracious readers that have swallowed almost every book, website, and conversation ever written. These machines, known as Large Language Models (LLMs), are the brains behind many of the chatbots and AI tools we use today. They learn by spotting patterns in words, trying to guess what comes next in a sentence. But here's the catch: if a computer has never heard a specific way of speaking, it gets confused. It might think a friendly joke is an insult, or it might fail to understand a cultural reference entirely. This is a big problem for "sociolects"—specialized languages used by specific groups of people, like teenagers, gamers, or the LGBTQ+ community. When AI doesn't understand these groups, it can accidentally be mean, misidentify people, or just fail to chat properly. The question researchers are asking is: How do we teach these digital brains to understand the rich, colorful, and evolving slang of queer communities without getting it wrong?

This is where a new project called SLAYING comes in. Think of it as a massive, community-approved dictionary and storybook specifically for queer slang. The researchers realized that while AI is getting better at understanding many types of language, it often stumbles over queer vernacular. Sometimes, the AI gets so confused that it mistakes a harmless, celebratory phrase for hate speech, or it generates a rude response to a user just because they used a specific word. To fix this, the team built the first real-world dataset of over 500 queer slang terms, complete with 3,405 examples of how these words are actually used in sentences. They didn't just pull these words from a textbook; they gathered them from movies, podcasts, and social media, and then had real members of the queer community check every single example to make sure it was used correctly and wasn't harmful.

The team used this new dataset to test some popular AI models, and they found something surprising. They discovered that a computer can be "nice" to queer people (meaning it doesn't hold stereotypes or hate them) but still be "clueless" about their language. It's like having a friend who loves you but doesn't understand your favorite inside jokes; they aren't being mean, they just don't get the reference. The study showed that when they taught an AI model using this new dataset of harmless, community-approved slang, the model actually got better at understanding queer people in general, not just the words.

However, the researchers also found that the AI isn't treating everyone equally. The models struggled the most with slang from specific groups, particularly those from African-American and Latine queer communities, as well as terms related to drag and non-binary identities. It's as if the AI has a blind spot for the very communities that created much of the slang it's trying to learn. Furthermore, the study showed that automated systems designed to catch "toxic" language often flag these queer terms as dangerous, even when they are being used in a positive, reclaimed way. This suggests that the very tools meant to keep the internet safe might be accidentally censoring queer voices. The paper suggests that by feeding these models more diverse, real-world examples of queer language, we can help them understand the community better, reducing both their confusion and their tendency to wrongly flag innocent speech as harmful.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →